amitdusane.com Adobe Analytics Learning

Start with the foundationsReport Suites

Privacy and Data Retention

An email arrives on a Monday morning. A customer who bought from you once, roughly two years ago, wants everything you hold about them erased. They are entitled to ask, the law gives you thirty days to answer, and the answer has to be true rather than comfortable.

Three small details in that request decide everything that follows, and none of them look important yet. It happened two years ago. They signed in to buy. They used a phone at the time, and have used a laptop ever since.

Each detail runs into a different setting in your report suite, and each of those settings was configured long ago by somebody who was thinking about conversion rates, not about this Monday. What follows is that single request, walked through to its answer.

Three gates, and the request has to pass all of them

Between the email and a truthful reply sit three questions. Does the data still exist. Can Adobe find the personal parts of it. Can the person in the email be matched to a visitor in the data.

Three gates, and a request has to pass all of them
Does it still exist? the retention policy set by contract, not by you Does Adobe know where? the privacy labels applied by hand, per suite Can you name them? the identity you collected a cookie, not an email A no at any gate ends the request and none of the three gates tells you it was the one that closed

The gates run in that order because Adobe enforces the first one. A data retention policy has to be in place before Adobe will process a privacy request at all, which makes sense once you see why: to promise that a person has been erased, there has to be an agreed edge to the data being searched. An archive with no edge cannot be promised about.

Take the request through one gate at a time.

Gate one: two years ago is closer to the edge than it sounds

Their purchase was twenty four months ago. Adobe's default retention policy is twenty five months.

They made it through by a single month. Had the email arrived next February instead, the honest answer would have been that nothing remains, and the whole exercise would have ended at the first gate.

Nobody at your organisation chose that margin. Twenty five months is simply where a company that has never discussed retention finds itself, and it is the number governing this request unless your contract says otherwise.

You can read the policy, but you cannot change it

The data governance dialog in the admin tools shows the period currently in force. Changing it is a conversation with your Adobe account team, not a setting. Reducing the period costs nothing. Extending it is a purchase made in yearly increments, eight at most, which puts the ceiling at ten years and one month. Administrators go looking for the dropdown. There is no dropdown.

That is worth a moment, because it contradicts everything else in this panel. Bot Filtering hands you a checkbox and steps back, on the reasoning that deciding which traffic counts is your judgment to defend. Retention takes the opposite position in the same screen.

The difference is what a mistake costs. An unticked bot rule leaves noise in a report, and the report survives to be corrected. A retention period set too short starts destroying data on a schedule, and no support ticket brings it back. A control whose worst outcome is total and irreversible does not belong behind a dropdown an administrator can reach on a Tuesday afternoon. The friction is the safety mechanism.

What the window does while nobody is watching

Retention is not a rule that fires once. It is a window that slides. The start is today minus your policy, the end is today, and every month the leading edge steps forward and takes another month of history with it.

The window does not sit still
Beyond the policy deleted, with no recovery Inside the policy reportable everywhere today older newer this edge moves forward every month, and takes history with it

What leaves is not moved to cheaper storage. Adobe reserves the right to delete it with no option for recovery, and that covers every surface at once: Analysis Workspace, Data Warehouse, and the feeds carrying raw hits into your own systems. One escape hatch exists, and only before the fact. A historical data dump can be requested ahead of a deletion, which turns data you are about to lose into a file you own.

Nothing announces the day your history stops

No notification, no banner, no entry in any log a practitioner reads. The first symptom is a year over year comparison returning nothing for the earlier period, and in a Workspace panel an absent year and a catastrophic one render identically. Both are zero. The person who discovers the retention boundary is usually the one presenting the trend, to an audience who assume the number means what it appears to mean.

The purchase in this request cleared that edge by four weeks. Nobody checked, and nobody would have been told.

Choosing the period, which is not an analytics decision

Left alone, an analyst sets retention to the maximum, on the reasonable grounds that data you no longer hold is data you can never study. Two other departments have standing to overrule that, and both are usually right.

Legal sets a floor and a ceiling at once. Some obligations require records to be kept. Purpose limitation, the principle under GDPR that personal data is held for a stated reason and no longer than that reason lasts, pushes the other way, and it reframes a long window entirely: an unnecessarily long retention period is not a prudent hedge, it is exposure you created and documented. Finance sets the price, because extensions are bought and a fourth year of history has to survive a renewal conversation.

Retention windowWhat it actually supportsWhat usually drives it
6 to 12 monthsCurrent trading only. No comparison against the same period last year is possible at any point in the calendarA strict purpose-limitation position, or cost
13 monthsOne year-on-year point: this month measured against the same month last year, and nothing widerThe practical floor for anyone reporting on seasonality
25 monthsA complete year compared against a complete prior year, with the current year still running underneath itAdobe's default, and the answer most organisations settle on
37 months and beyondMulti-year trend and long-horizon cohort workPurchased in yearly extensions, eight at most, to a ceiling of ten years and one month

Read that table as a one-way door. Every other setting in this module can be changed later at the cost of a seam in the data. Shortening retention changes what exists, and lengthening it afterwards does not bring back what the shorter window already discarded.

Record the comparison that justifies the window, next to the window

A period defended as "two years feels about right" loses to the first cost review, because a feeling does not survive a spreadsheet. A period defended as "the board pack compares each quarter against the same quarter last year, which needs twenty five months" survives, because it names something the business would visibly lose. Write down the comparison the window exists to protect, and keep it wherever the rest of the suite's decisions live.

One field never reaches gate one at all

The request asks for everything, which includes the IP address the purchase arrived from. For most report suites the honest answer is that there is not one, and there never was.

IP obfuscation sits in general account settings with three positions: leave the address alone, replace it with a hashed value, or remove it. Remove is selected by default on every newly created report suite, so unless somebody deliberately turned it off, this is already how your data behaves.

The usual belief about that setting is that it costs you something: privacy bought at the price of geographic reporting, bot rules, and the ability to exclude traffic by address. It is the intuition almost everyone arrives at, and the documented behaviour contradicts it.

Obfuscation happens last. IP filtering and exclusion, bot rule evaluation, and geo-segmentation lookups all finish before the address is touched. Country, region and city resolve normally. The exclude-by-IP tool described in Bot Filtering, the practical answer to automated traffic the IAB list never catches, keeps working exactly as before. Nothing in your reporting goes dark.

What you give up is narrower, and it is investigative rather than analytical.

You cannot exclude an address you can no longer see

Finding unwanted automated traffic usually means spotting a pattern in last quarter's data, identifying the address behind it, and excluding it. With the address removed at collection the middle step is gone: the pattern is visible and its origin is not. Exclusions can only be built from traffic arriving now. Like everything else in this panel the change is forward-only, so a suite switched to remove keeps whatever it already stored and never records another.

Gate two: Adobe has the data and still cannot find the person

The request clears the first gate. The purchase is inside the window, the hits are there, and the data has survived to be searched.

It is now going to fail anyway, for a reason that catches almost every organisation the first time.

Adobe Analytics knows the shape of your data completely and its meaning not at all. It knows eVar 12 exists, that it holds strings, and how many unique values arrived last month. Whether those strings are product categories or customer email addresses is knowledge held inside your company, sometimes inside one person, and never inside the platform.

Privacy labeling is how that knowledge gets written down in a form the Privacy Service can act on. It is the translation layer between your schema and a legal category, and until it exists a request has nothing to work with.

LabelWhat it declares about the fieldDependency
I1Identifies a person directly: a name, an email address, a phone numberPrerequisite for the ID and DEL labels
I2Identifies a person only in combination with something else: a loyalty number, a CRM identifierPrerequisite for the ID and DEL labels
S1Precise location, latitude and longitude accurate to roughly 100 metres or betterCan carry a delete label
S2Location broad enough to describe a geofenced area rather than a pointSensitive, but not a delete target by itself
ID-DEVICE and ID-PERSONThis field holds an identifier a request can be submitted against: a device, or an authenticated individualRequires I1 or I2
ACC-ALL and ACC-PERSONReturn this field's values in an access requestACC-PERSON needs an ID-PERSON field somewhere in the suite
DEL-DEVICE and DEL-PERSONAnonymise this field when a matching delete request runsRequires I1, I2 or S1

The dependency column is the part that bites. Identity labels are load-bearing: without I1 or I2 on a field, the delete label cannot be attached to it at all. Label the obvious personal fields as sensitive and stop there, and you have built a taxonomy that describes your data accurately while instructing the platform to do nothing.

Do this Label a variable for data governance
  1. Go to Analytics Admin All admin.
  2. Open Data configuration and collection Data governance, then select the report suite.
  3. Choose a variable group in the left filter: Standard Components, Conversion Variables (eVars), List Variables, Traffic Variables (props), Success Events or Classifications.
  4. Tick the variable, then open Edit Privacy Labels.
  5. Apply the labels, then choose Apply.

One variable group at a time, one report suite at a time. Labels belong to the suite they were applied in, exactly like bot rules and IP exclusions, which means the suite created next year for a new market starts with none of them.

The response that looks perfect and is wrong

Suppose the email address in this request sits in an eVar nobody labeled. Watch what the Privacy Service does.

It does not fail. It does not warn. It was told nothing about that field, so it looks at nothing, finds nothing, and reports that the work is complete. A small clean file goes back. The deletion is recorded as fulfilled. Legal replies inside the thirty days and the ticket closes.

Everyone in that chain believes the request was honoured. The customer's email address is still sitting in eVar 12.

The gap surfaces years later, and never at a convenient moment

An unlabeled variable produces no error at the time and no artefact afterwards, so the failure has no discovery mechanism of its own. It waits. What eventually surfaces it is an audit, a regulator's follow-up, or a second request from someone who has noticed their data is still in use, and by then the original response has been sent, logged and relied upon. Labeling is verified by opening the screen and reading it, because nothing in the system will ever raise a hand.

Labeling rests on an assumption worth examining, which is that you know which fields could be holding personal data. For eVar 12 you did. Somebody decided to put an email address there, and the only failure was never writing that decision down. The harder case is the field nobody decided anything about.

A page URL captured whole is that field. Query strings collect whatever the site, a partner, or a customer's own pasted link happens to put in them, and that includes the password reset token, the email address in a pre-filled form link, the session identifier appended by a third-party tool. Nobody chose to send any of it. It arrived because the implementation captures the URL, which is what every implementation does.

The variable nobody labeled because nobody knew it held anything

A labeling exercise walks the variable list and asks what each field is for, and the page URL answers "the page URL", which sounds like a location rather than a container. The personal data inside it is invisible to that question and stays unlabeled through every audit that asks it. The fix sits upstream of labeling: strip or allow-list query parameters at collection, so those values never reach a variable, rather than governing them once they have.

Gate three: they signed in, which is the only reason this works

The last detail in Monday's email is the one that saves it.

The request arrives as an email address, because that is how a person knows themselves. Adobe Analytics has never heard of them. It recognises the Experience Cloud ID, the legacy Analytics cookie that preceded it, and a custom visitor ID if your implementation was built to set one. An email address matches none of those.

So the request has to be translated first, and there are two routes. You can hold a person-level identifier in the data already, a CRM identifier in a custom visitor ID or a labeled eVar, and resolve the email to it using systems you control. Or you can collect the device identifiers straight from the browser while the person is asking, which is what Adobe's unified privacy JavaScript exists to do.

Now recall the phone and the laptop. The second route only reaches the browser in front of you, so it would find the laptop they use today and never touch the phone they bought on two years ago. Half the request would go unanswered, and the response would still come back clean.

Because they signed in, a person-level identifier was written on every one of those visits, from both devices. That single implementation decision, taken years earlier for reasons that had nothing to do with privacy, is what makes an honest answer possible.

Settle the identity route while nobody is waiting for an answer

This is the one part of privacy readiness that cannot be assembled after a request arrives, because the answer depends on data collected months or years earlier. Establish which identifier requests will be matched against, confirm it is being captured, and confirm it is labeled, as implementation work with no deadline attached. A statutory clock is a poor time to discover the identifier was never collected.

What the deletion then performs is worth knowing, because the word oversells it. Values are anonymised rather than removed. Traffic and commerce variables are overwritten with a random hexadecimal string prefixed "Data Privacy-", the visitor ID becomes a random value, custom visitor IDs and IP addresses are cleared, and precise coordinates are degraded to roughly a kilometre.

The hit itself survives all of it. Visits, page views and revenue do not move, and Tuesday's reports look exactly like Monday's. That is deliberate, and it resolves a tension that would otherwise be impossible: the individual becomes unfindable while the aggregate stays true. Honouring a person's rights does not require falsifying your own history.

All three fail quietly

One request, three gates, and every one of them decided years before the email arrived. Whether the purchase still existed came down to a contract nobody in analytics can edit. Whether Adobe could find the personal data came down to whether somebody once sat down and labeled eVar 12. Whether the person could be matched at all came down to a sign-in flow built for entirely different reasons.

The through-line is that all three fail quietly. History that ages out leaves no notice. A missing label returns a successful-looking response. An unmatched identity returns an empty result rather than an error. That is the same property running under every setting in this module, moved outside analytics into a domain where the questions come from regulators and your answer is already on the record.

All of this governs data you have already collected. Whether it should have been collected at all is the visitor's decision, expressed as a consent choice, and it has to be honoured before a single tag fires. Consent and Tag Management covers how a tag management system reads that choice and refuses to act without it.

Where to find it in Adobe Analytics

Privacy labeling: Admin > All admin > Data configuration and collection > Data governance, then select the report suite. The current retention policy is displayed here as well, for reading only; changing it goes through your Adobe account team.

IP obfuscation: Admin > Report Suites > Edit Settings > General > General Account Settings. Set per report suite, and forward-only, like the rest of that panel.

Need implementation steps?

This article focuses on the concepts, architecture, and practical guidance behind the topic. For the latest UI walkthroughs and step-by-step implementation instructions, use the links below. They leave this site and open Adobe's own documentation in a new tab.