Log in

The frontier labs have a data problem, and its solution is sitting in your corporate workspace.
Public knowledge has, by and large, already been consumed by the AI labs. They have moved on to buying human-generated data at scale, and each of the frontier labs is now reported to spend on the order of a billion dollars a year doing it. Some of this is bought from providers whose users generate the useful AI signal: Reddit and StackOverflow signed licensing deals with OpenAI and Google worth tens of millions a year each. Some of it comes from companies that have shuttered. SimpleClosure launched an Asset Hub for winding-down startups to sell their code, Slack history and email, and reports that it has processed around 100 such deals in a year. Forbes reported one defunct company selling 13 years of Slack, email, and Jira tickets for several hundred thousand dollars.
What the labs want next is the thing every company has and none has yet sold: corporate operational data. Not the product, not the customer data itself or the plumbing. Email, chat, calendar, files, tickets, the HRIS, the CRM, the contracts, the decks. It's the record of how work actually happens, which is exactly what makes it valuable for post-training agents to perform business tasks, and also what makes it dangerous to move. This is a new corner of AI data licensing, and the rules for it haven’t been decided.
Corporate operational data contains the data on every person who ever worked at the company and a good number who bought from it, too.
A corpus of workplace systems is a corpus of people. One employee shows up in a hiring thread, then a payroll record, then eight years of messages, then a performance review, then an exit note, under a name, a nickname, an email alias and an employee ID that don't obviously connect to each other. Customers appear in support tickets with their addresses and their complaints. Candidates who were never hired are there too. So are the salary numbers, the medical leave requests, and the thing somebody said about their divorce in a direct message in 2019.
Data anonymization, or de-identification, is the work of removing that identity while leaving the patterns intact. Resolving one person to one synthetic identity across all of it (aka maintaining referential integrity) is the part that separates real de-identification from a find-and-replace. The goal is a corpus that still reads like a company, with synthetic people in place of real ones, because a lab that buys an unreadable corpus has bought nothing. Preserving the AI signal is equally as important as detecting the sensitive entities across 100s of terabytes of data and dozens of internal systems.
The buyer is training models on the data it acquires, which means that data risks resurfacing in production.
Post-training writes what it sees into model weights. Fine-tuning is worse than pre-training on this count. The datasets are smaller and more valuable, and the shift in the model's own probabilities becomes a signal an attacker can use to prove a particular record was in the training set (arXiv 2602.00688). A 2025 IJCAI survey of PII leakage in language models found models reproducing training sequences verbatim, personal information included, through nothing more exotic than the public API.
The attacks that do this are not sophisticated. In an email-extraction attack, someone prompts the model with the fragments they already have, like a name and an employer, then asks it to complete the rest and checks whether what came back is real. Measurements published this year put average success rates for that kind of attack around 78%.
And cleaning up afterwards doesn't work. Redacting a name in the source doesn't help once context-based attacks can reconstruct it, and unlearning at the parameter level is still a research problem rather than something you can buy. A file can be deleted. A weight can't. Whatever identity survives into post-training stays in the artifact for the life of the model, and it leaves through the same API the lab's customers are paying for.
This is the part most sellers have backwards.
Under the CCPA, the obligations sit with the business that collected the information. Under HIPAA, the covered entity stays responsible even when a business associate agreement pushes the work downstream. Under GDPR, the controller is on the hook for it, and the controller is usually the company that gathered the data and still holds it, which is to say the company trying to monetize it. When Toysmart tried to sell its customer database out of bankruptcy in 2000, the FTC sued over Toysmart's own privacy promise. The list of roughly 190,000 customers was ultimately bought for $50,000 by a Disney subsidiary that was also a creditor, on the condition that it be destroyed rather than transferred, with an affidavit filed with the court to confirm it. Realized value: zero. Fifteen years later, 38 state attorneys general and the FTC objected to RadioShack's plan to sell its customer data, and the settlement let the estate transfer only email addresses collected in the prior two years plus limited transaction data, with an opt-out for customers and a bar on resale, while roughly 50 million customer files were destroyed. No card numbers, Social Security numbers, dates of birth or phone numbers moved at all.
You can outsource the work but the exposure still lives with the seller. And there's a second-order problem: whether a corpus even counts as de-identified isn't a fixed property of the file. It depends on who is holding it and what they could still do with it, which means the identity and competence of whoever performed the de-identification is part of the legal answer rather than a procurement detail.
Three candidates. Only one secure answer.
Option 1: the seller. They have the liability, which is the right incentive, but they usually have neither the tooling to de-identify company data at scale, across a hundred million records and a dozen systems, nor any standing to certify their own work. A determination you write about yourself is not usually a credible determination.
Option 2: the buyer, or the broker between the seller and buyer. They often have real capability. But they also have a structural conflict, because they are paid to move the corpus rather than to hold it. If the broker de-identifies, the broker funds the work and captures the savings from doing less of it. There's no attestation, no sampling record, no independent party who looked at the output and signed their name to it. If recall came in at 92% rather than 97%, the corpus still ships, the deal still closes, and the missing 8% surfaces years later inside a model that millions of people are querying.
None of this requires anyone that’s behaving badly. It simply requires a company that is optimizing the thing it's measured on, which in the middle of this market is speed and growth trajectory.
Option 3: an independent third party, retained by the seller. Within the business transaction, this is the only participant whose product is the de-identification and attestation itself, and therefore the only one whose incentive points at getting it right rather than getting it done.
Health care worked this out decades ago. HIPAA lets you de-identify by expert determination which is a named person with statistical training that applies accepted methods, concludes the re-identification risk is low, and documents how they got there. They provide a certification that lasts for 1-2 years.
Court-supervised sales work the same way. Federal bankruptcy law won't let a trustee sell personal information in conflict with the company's own privacy policy unless a consumer privacy ombudsman is appointed and the court approves the sale after a hearing. An independent party, on the record, is required before the data moves.
Every regime that has thought hard about selling personal information has reached for independence. Corporate operational data is the one place where the stakes are comparable and the practice is absent, where the standard is a reasonable-measures test with no expert, no methodology and no documentation required, and where two companies can claim the same compliance while one resolved every employee's identity across every system and the other ran a find-and-replace on a spreadsheet of names.
I should say plainly that I help run a company that does this work, so discount the conclusion accordingly. The argument isn't mine, though. It belongs to HIPAA, to the bankruptcy code, and to every regulator who has looked at a data sale and asked who was watching.
To learn more about de-identifying company data and how Tonic.ai can help, reach out to us or book time directly with our team.

Tomer Benami is the VP of Finance and Bizops at Tonic.ai where he brings a blend of core finance expertise, operational savvy, and vision to go-to-market activities. With a proven track record of serving as the senior-most finance leader at companies such as VirtualHealth and Apploi, Tomer enjoys partnering with executive teams, steering organizations towards strategic goals and delivering meaningful results. Beginning his career at KPMG and holding a Master's Degree from the University of Washington, Foster School of Business, he is enthusiastic about the transformative potential of AI while advocating for its responsible and ethical utilization in shaping our future.