Data privacy

Our core tenets for de-identification in AI data licensing

Author
Ian Coe
September 23, 2026

Data has always changed hands. What has shifted is what it's worth and to whom.

The first large AI models learned from the public internet. That source is finite, and the models have grown faster than it has. What remains scarce is the record of how work actually gets done: the support tickets, the project files, the internal correspondence where decisions were made. That record exists almost nowhere in public. It exists inside companies.

So companies are now selling and licensing their operating data to train AI models. These transactions are large, they are accelerating, and no standard governs how the people inside those datasets are protected when one happens.

That gap is not a decision anyone made. It is the ordinary lag between a practice and the rules that eventually govern it, and right now the practice is well ahead.

Some of this data is consumer data, and consumer data has at least a partial framework around it. GDPR in Europe, CCPA in California. Much of it is employee data, and employee data sits outside most US state privacy laws by explicit exemption. But the more important point is narrower than either: no framework anywhere addresses this kind of transfer directly. Not the volume of it, not the fact that the buyer is training a model, and not the position of a person who is in the dataset because they worked somewhere, and who was never offered a choice about it.

We do not think the answer is to wait. Regulation will come, and it will be better written for having a real practice to describe. But the practice is being set now. The terms accepted in the first transactions become the terms buyers expect in later deals.

The absence of a rule is not the absence of an obligation. Standards written now, ahead of regulation, have to be worth adopting. They have to come from people who have done the work, because the ways this goes wrong are specific and mostly invisible from the outside. And they have to hold up for everyone the transaction touches: the people in the data first, but also the seller who needs the transfer to survive scrutiny, and the buyer who needs to know what it is actually receiving.

We are paid to do this work. That is precisely why we should be explicit about the conditions under which we will do it.

Transfers are one part of our work. Companies also use Tonic.ai to de-identify data they keep: test data for their own engineers, training data for their own models. Some of what follows applies to all of it. Some of it exists only because there is a buyer. When data leaves the company that collected it, a new party enters with interests of its own, and the people in the data end up in the hands of an organization they never dealt with. That is a narrower problem than de-identification in general, and it fails in its own ways.

These are the seven rules Tonic.ai works by when data we de-identify is sold or licensed, and the rules we believe any such transfer should meet. Each one ends with how it can be checked, because a rule that can't be checked provides no real assurance to anyone.

1. Retain a de-identification provider with no stake in the sale

The de-identification provider must be retained by the seller, take direction from the seller and its counsel, and be paid for work performed whether or not the sale closes.

The buyer holds no stake in the provider and does not direct what is transformed or removed. Any commercial relationship between the provider and any party to the transaction, the buyer included, is disclosed.

The payment structure is key. A provider paid on completion of the sale has an interest in completion of the sale. A provider paid for the data de-identification work has an interest in the data de-identification work. This distinction determines everything downstream, because every judgment call in a de-identification project is a decision about whether to spend more effort or accept more risk. The people in the data need those calls made by someone who is paid to meet the standard, not by someone who is paid more if the transaction completes. The seller and the buyer need an assessment that neither of them shaped, because an assessment either party influenced will not hold up when a court, a regulator, or an objecting party examines it.

Verification method: The engagement documents state who retains the provider, who directs the work, and how the provider is compensated, and where a court is involved those terms are placed on the record in a sworn declaration.

2. Get the buyer's commitment in writing before the data moves

Transforming the data is one layer of protection. The rest depends on what the recipient is permitted to do with it afterward.

Before transfer, the buyer must commit in writing to hold the data for internal, access-controlled use, and not to attempt to re-identify anyone in it.

This matters because no method removes all risk, and the written commitment is what covers what the method cannot. A transaction that transforms the data thoroughly but leaves the recipient unconstrained has protected the file rather than the people in it. For the buyer, the commitment is also an asset: a documented limit it can demonstrate it operated within.

Verification method: The sale documents contain the recipient's written use limitations and its undertaking not to re-identify, available for review by the parties and any court-appointed overseer.

3. Never deliver the source data or the mapping

The buyer receives the processed output only. Never the source data, never a pre-transformation copy, and never the record that connects real values to their replacements.

That mapping stays inside the provider's access-controlled environment, encrypted and audit-logged. It is destroyed when the engagement ends, on terms the seller directs, along with the provider's copy of the source data, with written confirmation.

This is the rule the others rest on. A recipient who holds the mapping has not received de-identified data; it holds the original data in a form it can reverse at will. Every protection described elsewhere in this list assumes the mapping never transferred to the buyer.

Verification method: Written destruction confirmations record the disposal of the mapping and the provider's copy of the source data, and the processing environment operates under controls covered by a SOC 2 Type II report validated by an independent auditor.

4. Assume the people in the data can be looked up

Privacy frameworks tend to imagine a large and anonymous population. A company's operating data describes a specific and finite group of people, and often a group whose membership is a matter of public record.

Consider what it takes to check a name. A former workforce is enumerable. Professional profiles are public and searchable. Anyone curious enough to try can assemble a list of who was there and compare it against a dataset. This is not a sophisticated adversary with unusual resources. It is one person with a browser, and it is a legitimate concern for the people in the data.

The risk analysis must begin from that assumption rather than from an average case. Workforce identifiers, employee numbers among them, are classified as personal information at the outset rather than discovered midway through, alongside the indirect identifiers that can single out a person in combination even when no direct identifier remains. Removing obvious names is not sufficient, and a method that would satisfy a casual observer is not the standard.

Verification method: A written methodology names the adversary the analysis assumes and the controls relied on against that adversary, and is shared with the seller and any court-appointed overseer before processing begins.

5. Measure the result against thresholds set in advance

A statement that the data was cleaned is an assertion. A measurement is evidence.

Accuracy must be measured on human-labeled samples, with precision and recall calculated for each category of personal information, and reported per individual as well as per mention. A single missed mention can expose a person who appears throughout the corpus. Each round draws a fresh random sample, so performance reflects the method rather than familiarity with one set of records. The finished output is sampled again before delivery.

The thresholds and the reasoning behind them are written down and shared with the seller, its counsel, and any overseer before the full corpus is processed. Thresholds set before the work begins are thresholds the work has to meet. Thresholds chosen after the results arrive tend to land wherever the results already are. The seller should be able to hand the samples and the results to a reviewer of its own choosing, and have that reviewer reach the same finding about whether the thresholds were met.

Verification method: A methodology report states the accuracy thresholds, the sampling approach, and the measured results for each category of personal information, in a form the seller can pass to an independent reviewer.

6. Exclude what cannot be made safe

Some content resists de-identification. A scan too degraded to read reliably. A file format the tooling does not handle. A record where the identity of a person cannot be separated from the substance of what is being described without destroying both.

An additional category sits alongside this. A company's mailboxes and drives hold material belonging to its suppliers and partners. Replacing names does nothing for a pricing schedule or an engineering drawing, because what makes those sensitive is confidentiality rather than identity. Material like that is removed as whole documents, under criteria the seller's counsel sets, rather than edited.

In every one of these cases the content is excluded from delivery or escalated to counsel, and the exclusion is logged. Nothing is delivered on the assumption that it is probably fine. We have left material out of deliveries under this rule more than once, and we would be skeptical of any provider claiming it never had to.

Verification method: The material received at the start of the engagement is counted, and that count is reconciled against what is delivered at the end. An exclusion manifest traces every removal to the rule that required it, so the parties can demonstrate that nothing was lost or added beyond the documented exclusions.

7. Sign the certification, and state what it does not cover

At corporate scale, some personal information will get past any detection method. A reasonableness standard requires saying so, and saying what the remaining risk depends on.

A named individual able to defend the work must sign the certification. The certification, together with the methodology report it accompanies, states the standard applied, the method, the measured results, and the residual risk, in writing. That person must be available to explain it to the court and to any overseer, under oath if called.

The six rules above describe how the work should be done; this rule puts a name on whether it was. The people in the data learn what was done and what was not. The seller and the buyer receive a standard they can hold a person to, rather than a reassurance they have no way to examine.

A certification reporting no residual risk at all is the one we would trust least. At corporate scale, some residual risk always exists. A document that does not mention it is not telling you the risk is zero. It is telling you that nobody measured it.

Verification method: A signed certification, read alongside the methodology report delivered with it, states the standard, the method, the measured results, and the residual risk, attributed to a named individual available for examination.

The standard we want to be held to

The technical work here is genuinely hard. Detecting personal information across millions of records, in formats that actively resist it, is a problem we have worked on since 2018, and one that keeps changing as the data grows in volume and variety.

But the seven rules above are not technical achievements. They are decisions, made in advance and in writing: what happens when a record cannot be processed safely, who is permitted to hold the mapping, what gets written down when a measurement comes back lower than anyone hoped. Any organization willing to settle those questions before the work is underway and on a deadline can meet this standard.

These transactions are going to continue, and the concerns people have raised about them are legitimate. Each transaction completed without a standard becomes a precedent for the next. The terms of this market are being written in its first deals, whether or not anyone writes them down.

Handled under the standard we have outlined here, corporate data can move without the people inside it bearing the cost.

We expect regulation here, and we would welcome it. When it arrives, we hope it is written by people who understand where de-identification actually fails, and written with the people in the data as its first concern.

Until then, these are the rules we hold ourselves to. We would be glad to see them become the rules everyone is held to.

Ian Coe
Co-Founder & CEO

Ian Coe is the Co-founder and CEO of Tonic.ai. For over a decade, Ian has worked to advance data-driven solutions by removing barriers impeding teams from answering their most important questions. As an early member of the commercial division at Palantir Technologies, he led teams solving data integration chal- lenges in industries ranging from financial services to the media. At Tableau, he continued to focus on analytics, driving the vision around statistics and the calculation language. As a founder of Tonic, Ian is directing his energies toward synthetic data generation to maximize data utility, protect customer privacy, and drive developer efficiency.