Log in
PrivacyBench v2 scores de-identification end to end: raw email, Slack, and document exports in, a coherent synthetic export out. See how Tonic Textual compares with six frontier LLM pipelines.

PrivacyBench is Tonic.ai's open benchmark for de-identifying workspace data, part of the Tonic.ai AI model benchmark library. Version 2 scores the whole job end to end. A pipeline receives raw .eml emails, a Slack workspace export, and PDF, Word, Excel, and CSV documents, then has to detect the PII, replace it with coherent synthetic values, and write the export back out in its original file formats. Across seven pipelines, Textual Graph Synthesis, the workspace de-identification pipeline built on Tonic Textual, scored highest on every metric and reached 86.2% on the combined detection-and-synthesis score, 6.1 points ahead of the best pure-LLM pipeline. Textual also lifted the LLM it ran with, GPT-5.6 Sol, from 73.4% on its own to 86.2%, a 12.8-point gain.
21 workspace exports · 18,040 messages · 357 documents (PDF, DOCX, XLSX, CSV) · 10 entity types · 132,909 labeled PII spans · 7 pipelines tested · human-validated answer key
If you've seen the PrivacyBench v1 results, v2 is a harder test, and its scores aren't comparable with v1's. See what changed from PrivacyBench v1 and the archived v1 results further down this page.
Teams de-identify their email, Slack, and shared-drive exports so they can put that data to work: training and fine-tuning models, building agent reinforcement learning environments (RLE), licensing operational data to AI labs, and sharing data with partners. None of that works if real PII survives in the output, and none of it works if the replacements break the data.
That data never arrives as a tidy list of sentences. It arrives as an export: a mix of semi-structured files and free text. One person shows up as a display name in an email header, a handle in a Slack message, a name in the signature block of a Word document, and a row in a spreadsheet. A de-identification pipeline has to find every one of those mentions and map all of them onto a single synthetic identity, in every file. PrivacyBench v2 scores the whole export together, so a pipeline that renames someone one way in the emails and another way in the documents is scored as incoherent.
Every pipeline has to complete four stages:
A strong replacement has to hold four things together:
Existing NER benchmarks, like CoNLL, OntoNotes, and TAB, score detection well. None of them score the quality of the replacement, and none of them score it across a real export. That gap is what PrivacyBench fills.

A benchmark result only helps you if the test resembles the job you need done and you can check the result yourself. PrivacyBench is built on both.
Tonic.ai built this benchmark and the system that scores highest on it. That's the reason every piece of it is public: so you can score your own pipeline against the same data and compare directly with the results here. None of Tonic.ai’s models were trained on any of the data in PrivacyBench.
Every PrivacyBench dataset is fully synthetic. Tonic Fabricate generated each one from a seed that declares every character's PII before any text exists, which makes the ground truth a byproduct of generation rather than a hand-labeling project. That solves the two problems that make PII benchmarks expensive: a public benchmark can't contain real PII, and trustworthy labels are slow and costly to tag by hand.
Each of the 21 datasets is one fictional protagonist's workspace, set in a different industry, including pharma, retail, finance, airlines, tech, and manufacturing. Together they hold 18,040 messages (2,540 emails and 15,500 Slack messages) and 357 supporting documents, 17 per dataset: 10 PDFs, 3 Word documents, 2 Excel workbooks, and 2 CSVs. In 11 datasets the documents travel as email attachments and Slack file shares; in the other 10 they live in a Drive folder.
The answer key covers 132,909 PII spans across ten entity types: given name, family name, email address, username, organization, phone number, location address, employee ID, account number, and URL. Each span is tied to the character or organization it belongs to, which is what lets the benchmark check that one real identity maps to one synthetic identity.
PrivacyBench reports three scores: one for detection, one for the quality of the replacements, and one that combines both to score the entire pipeline. The definitions are unchanged from v1.
Each of the first two scores hides something on its own. Recall is a privacy measure: it tells you how much PII was caught, but nothing about whether the de-identified data is still usable. Synthesis accuracy is a utility measure, and on its own it can be gamed: detect a handful of easy entities, replace them perfectly, and post a high number while the hard PII slips through. The combined score captures privacy and utility together, which is why it's the number to rank by.

Tonic Textual's Graph Synthesis pipeline led all seven pipelines on every metric. It pairs Textual's NER with graph-based synthesis that links every mention of a person or organization across the export and replaces them consistently. The six other pipelines are pure-LLM agents: each model wrote its own extraction and rendering code in a sandbox, then performed NER and synthesis itself.
PrivacyBench v2 results: macro average across 21 datasets, ordered by the combined score
Two results explain Textual's lead.
Textual's NER model is a small, efficient encoder-based model, not an LLM. That makes it much faster and cheaper to run than asking an LLM to find the PII, and on PrivacyBench it still detects PII better than the frontier models do. End to end, it caught slightly more of the known PII than any pure-LLM pipeline, and against the human annotations it posted the highest F1 of any engine tested, as the detection results below show.
The even more definitive lead is in synthesis. Finding a person's name is one job. Giving that person the same new identity in every email, Slack message, and document they appear in is a harder one, and it's where the pure-LLM pipelines lose ground. Textual's graph-based synthesis links every mention of a person or organization across the export before it replaces anything, so its replacements stayed coherent more often than any other pipeline's: 91.8% of the time, against 87.5% for the next best.
Put those two together and you get the clearest result in the benchmark. Textual Graph Synthesis used GPT-5.6 Sol for its LLM steps, which makes the GPT-5.6 Sol row a direct comparison: the same model, with and without Textual. Running the whole job on its own, Sol scored 73.4% and finished third among the pure-LLM pipelines, behind GPT-6 Astra and Claude Opus 5. Inside Textual, the same model scored 86.2%, a 12.8-point gain that took it from third place to first. Nothing about the model changed; what changed was the system around it.

The end-to-end results measure recall against the generated answer key, which covers the seed characters and their organizations but not every name that appears in the data. To score detection completely, human annotators labeled every PII span in the emails, Slack messages, PDFs, and Word documents of 6 of the 21 datasets: 25,918 spans across 5,198 messages and 656 document pages. Against that human gold, precision and F1 become meaningful.
NER scores against the human annotations (overlap matching)
Tonic Textual posted the highest F1 of any engine, 91.7%, and it's the only engine above 90% on precision, recall, and F1 at once. Some LLMs scored up to 2 points higher on precision, but they gave up more than 8 points of recall to get there. For privacy, recall is the number that matters: a precision miss over-redacts something that wasn't PII, while a recall miss leaves real PII in your data. The two open-source engines trailed further, with Presidio at 82.8% F1 and GLiNER2 at 75.1%.
The Textual NER models were not trained on any PrivacyBench data, and the Textual via SDK row is reproducible with a Textual API key and the open scoring code.

PrivacyBench v1 scored email and Slack messages supplied as pre-extracted text, across five entity types, with two-stage pipelines that ran NER and then synthesis. v2 keeps the v1 email and Slack messages, now in their raw export formats, along with the same three metric definitions, and turns the task into a full de-identification job:
Because the task is harder, v2 scores run lower than v1 scores, and the two versions aren't comparable. Comparing them is like comparing a 5K time to a trail-race time: both measure how fast you are, on very different ground. The v1 data and scoring code remain available at the v1 revision of the dataset and the v1 tag of the metrics repo. We've also re-run v1 on the current version of Textual, so the v1 results below reflect how Textual performs today.
We've re-run PrivacyBench v1 using Textual Graph Synthesis, which is used for PrivacyBench v2. At the time of publishing v1, Textual Graph Synthesis was still in development.
On v1, Textual Graph Synthesis reached 95.4% on the combined synthesis + NER score, with 98.2% NER recall. The original v1 pipelines appear below as first published, alongside the new run.
PrivacyBench v1 results: the current Textual run and the original v1 pipelines
The Opus-only pipeline still posts the highest raw synthesis accuracy on v1, 99.0%, but only because its lower recall left it replacing the easier entities it caught. With detection misses counted, its combined score is 87.7%.

Human annotators fully labeled the email and Slack messages in 6 of the 21 datasets. That makes it possible to score v1 detection on precision, recall, and F1, rather than recall alone. These are the same 6 datasets used in v2.
NER scores against the v1 human annotations for overlap metric
Textual is the clear winner for overall NER when precision, recall, and F1 are all considered across the six datasets. For both scoring methods, Textual's precision, recall, and F1 are all over 90%. It is the only model that achieves over 90% for all three scores, and it's the only model that achieves over 90% recall, which it does for both scoring methods. There are other models that score similarly to Textual on precision, but Textual wins by a lot on recall in those cases. Recall is the more important metric when considering privacy as it measures privacy leaks.
If you're de-identifying workspace data for AI, PrivacyBench v2 points one way: frontier LLMs working alone trail on keeping identities coherent across a full export, and Tonic Textual's graph-based synthesis closes that gap. The combined score, the same-model gain, and the detection results against human annotations all point the same direction.
The strongest case is the one you run yourself. Point your pipeline at the 21 public workspace exports, score it with the open metrics code, and compare your numbers with the ones here. To reproduce the Textual NER results, create a free Textual account for an API key and follow the instructions on the dataset card.
Tonic Textual is also the engine behind Tonic.ai's company data de-identification work, which prepares operational data like email, chat, tickets, and documents for licensing and model training. For the principles behind that work, read the core tenets of de-identification for AI data licensing. To see Textual on your own data, book a demo.
v2 turns PrivacyBench into a full de-identification job. Pipelines now receive raw export files instead of extracted text, handle PDF, Word, Excel, and CSV documents alongside email and Slack messages, detect ten entity types instead of five, and write the synthetic export back out in its original formats. v2 also adds human-annotated ground truth for six datasets.
v2 turns PrivacyBench into a full de-identification job. Pipelines now receive raw export files instead of extracted text, handle PDF, Word, Excel, and CSV documents alongside email and Slack messages, detect ten entity types instead of five, and write the synthetic export back out in its original formats. v2 also adds human-annotated ground truth for six datasets.
No. v2 is a harder task with different data, so scores drop across the board and the two versions aren't comparable. Compare pipelines within a version, not across versions.
Yes. Against the human annotations, Tonic Textual posted the highest F1 of any engine tested, 91.7%, ahead of every frontier LLM and both open-source engines. Some LLMs scored slightly higher on precision but more than 8 points lower on recall, and recall is what determines whether real PII stays in your data.
Synthesis accuracy measures replacement quality: of the PII a pipeline detected and replaced, how much is coherent? An LLM judge checks each replacement, per character, against the synthetic identity it should belong to, so a nickname has to map to a nickname form of the same synthetic name, and a synthetic email address has to belong to the synthetic person.
Yes. The 21 workspace exports and their answer keys are public on Hugging Face, and the scoring code is on GitHub. Run your pipeline on the raw exports, write its replacements in the benchmark's output format, and score them against the same ground truth and metrics used here.
A public PII benchmark can't contain real PII, and trustworthy labels are expensive to annotate by hand. Because Tonic Fabricate generates each dataset from a seed that fixes every character's PII in advance, the ground truth comes with the data, and the difficulty can be dialed up on purpose. Human annotators then validated that answer key on six of the datasets.

Joe Ferrara is Head of AI at Tonic.ai, where he uses the latest developments in artificial intelligence to improve named entity recognition and synthetic data generation in Tonic Textual. He holds a Ph.D. in Mathematics from the University of California, Santa Cruz, and a B.A. in Mathematics from the University of California, Berkeley. Prior to his role at Tonic.ai, Joe served as a Data Scientist at ICW Group, where he gained experience applying traditional data science techniques in the context of the insurance industry. His background in theoretical mathematics research gives him a unique perspective on his work in artificial intelligence and machine learning.