AI model benchmarks

Why Tonic Textual beats frontier LLMs at end-to-end de-identification: PrivacyBench v2 results

PrivacyBench v2 scores de-identification end to end: raw email, Slack, and document exports in, a coherent synthetic export out. See how Tonic Textual compares with six frontier LLM pipelines.

July 8, 2026
Last updated
September 29, 2026
0 min read

PrivacyBench is Tonic.ai's open benchmark for de-identifying workspace data, part of the Tonic.ai AI model benchmark library. Version 2 scores the whole job end to end. A pipeline receives raw .eml emails, a Slack workspace export, and PDF, Word, Excel, and CSV documents, then has to detect the PII, replace it with coherent synthetic values, and write the export back out in its original file formats. Across seven pipelines, Textual Graph Synthesis, the workspace de-identification pipeline built on Tonic Textual, scored highest on every metric and reached 86.2% on the combined detection-and-synthesis score, 6.1 points ahead of the best pure-LLM pipeline. Textual also lifted the LLM it ran with, GPT-5.6 Sol, from 73.4% on its own to 86.2%, a 12.8-point gain.

21 workspace exports · 18,040 messages · 357 documents (PDF, DOCX, XLSX, CSV) · 10 entity types · 132,909 labeled PII spans · 7 pipelines tested · human-validated answer key

If you've seen the PrivacyBench v1 results, v2 is a harder test, and its scores aren't comparable with v1's. See what changed from PrivacyBench v1 and the archived v1 results further down this page.

What PrivacyBench measures, and why v2 goes end to end

Teams de-identify their email, Slack, and shared-drive exports so they can put that data to work: training and fine-tuning models, building agent reinforcement learning environments (RLE), licensing operational data to AI labs, and sharing data with partners. None of that works if real PII survives in the output, and none of it works if the replacements break the data.

That data never arrives as a tidy list of sentences. It arrives as an export: a mix of semi-structured files and free text. One person shows up as a display name in an email header, a handle in a Slack message, a name in the signature block of a Word document, and a row in a spreadsheet. A de-identification pipeline has to find every one of those mentions and map all of them onto a single synthetic identity, in every file. PrivacyBench v2 scores the whole export together, so a pipeline that renames someone one way in the emails and another way in the documents is scored as incoherent.

Every pipeline has to complete four stages:

  1. Extract the text from raw export files, including base64-encoded email attachments.
  2. Detect the PII with named entity recognition (NER).
  3. Synthesize a coherent replacement for every detected entity.
  4. Render the synthetic export back out in the same file formats as the input.

A strong replacement has to hold four things together:

  • Cross-message consistency: every mention of one person maps to the same synthetic identity across every message.
  • Cross-file consistency: that same identity holds in the PDFs, Word documents, spreadsheets, and CSVs, not just in the messages.
  • Within-identity derivation: the synthetic email address, handle, phone number, and employee ID plausibly belong to the synthetic name, and a nickname maps to a nickname form of it.
  • Organization coherence: every surface form of one employer maps to a single synthetic organization, with dependent values like email domains, addresses, and URLs kept in step.

Existing NER benchmarks, like CoNLL, OntoNotes, and TAB, score detection well. None of them score the quality of the replacement, and none of them score it across a real export. That gap is what PrivacyBench fills.

Tonic Textual detects entities, links every mention to one individual, synthesizes consistently, and preserves each entity's type and format across email, Slack, and documents.
Textual detects entities, links every mention to one individual, and replaces them consistently across every source, the same coherence PrivacyBench scores.

What makes a de-identification benchmark worth trusting

A benchmark result only helps you if the test resembles the job you need done and you can check the result yourself. PrivacyBench is built on both.

  • Real export formats, not flattened tables. Inputs are RFC 5322 .eml files with base64 attachments, a Slack workspace export, and a Google Drive-style file listing. Every labeled span is located in its file's native coordinates: byte offsets in an email, JSON pointers in a Slack day file, per-character boxes on a PDF page, cells in a spreadsheet.
  • Realistic narrative, not template fill. Tonic Fabricate generated a month of email and Slack activity for each of the 21 fictional protagonists, keeping one coherent storyline across the whole period. The supporting documents were generated from those storylines and are referenced in the conversations the way real work files are.
  • Linked and realistic at once. Every identity is linked across files for scoring, the unstructured-data equivalent of referential integrity, and the data still reads like a company doing its work: the kind of messy, cross-referenced text your own exports contain.
  • Deliberately hard attribution. The seeds include characters who share a name, characters with nicknames, and messages where only the surrounding conversation tells you which person is meant.
  • Scores that separate systems. Each of the three metrics spreads the pipelines tested: NER recall ranges from 54.7% to 93.9%, and synthesis accuracy from 77.4% to 91.8%. No metric is one that every system passes by construction.
  • An answer key humans checked. Annotators reviewed the generated labels for 6 of the 21 datasets and confirmed 95.8% of them exactly, then added the spans the generator missed so detection could be scored on precision, recall, and F1.
  • Current competition. The end-to-end comparison covers six agent pipelines built on frontier models from OpenAI and Anthropic, and the detection comparison adds two widely used open-source NER models, Microsoft Presidio and GLiNER2.
  • Open data, open answer key, open scoring code. The full dataset and ground truth are public on Hugging Face, the scoring code is open source on GitHub, and the Textual NER scores are reproducible with the Textual SDK.

Tonic.ai built this benchmark and the system that scores highest on it. That's the reason every piece of it is public: so you can score your own pipeline against the same data and compare directly with the results here. None of Tonic.ai’s models were trained on any of the data in PrivacyBench.

How the benchmark data was built

Every PrivacyBench dataset is fully synthetic. Tonic Fabricate generated each one from a seed that declares every character's PII before any text exists, which makes the ground truth a byproduct of generation rather than a hand-labeling project. That solves the two problems that make PII benchmarks expensive: a public benchmark can't contain real PII, and trustworthy labels are slow and costly to tag by hand.

Each of the 21 datasets is one fictional protagonist's workspace, set in a different industry, including pharma, retail, finance, airlines, tech, and manufacturing. Together they hold 18,040 messages (2,540 emails and 15,500 Slack messages) and 357 supporting documents, 17 per dataset: 10 PDFs, 3 Word documents, 2 Excel workbooks, and 2 CSVs. In 11 datasets the documents travel as email attachments and Slack file shares; in the other 10 they live in a Drive folder.

The answer key covers 132,909 PII spans across ten entity types: given name, family name, email address, username, organization, phone number, location address, employee ID, account number, and URL. Each span is tied to the character or organization it belongs to, which is what lets the benchmark check that one real identity maps to one synthetic identity.

How PrivacyBench scores de-identification: three metrics

PrivacyBench reports three scores: one for detection, one for the quality of the replacements, and one that combines both to score the entire pipeline. The definitions are unchanged from v1.

Privacy-bench metrics: what each one measures
Metric What it measures
NER recall Of the PII known to be in the data, how much the pipeline detected. A miss here is a privacy leak: real PII left in the output.
Synthesis accuracy Of the entities the pipeline detected, how many it replaced coherently. An LLM judge scores each replacement, per character, against the synthetic identity it should belong to.
Synthesis + NER accuracy The bottom-line score. It counts detection misses as failures, so a pipeline can’t inflate its synthesis score by detecting only the easy entities.

Each of the first two scores hides something on its own. Recall is a privacy measure: it tells you how much PII was caught, but nothing about whether the de-identified data is still usable. Synthesis accuracy is a utility measure, and on its own it can be gamed: detect a handful of easy entities, replace them perfectly, and post a high number while the hard PII slips through. The combined score captures privacy and utility together, which is why it's the number to rank by.

One message scored end to end: detection misses and incoherent replacements both pull the combined score below either stage on its own.
One message scored end to end. Detection misses and incoherent replacements both pull the combined score below either stage on its own. v2 computes the same metrics across entire exports.

What the benchmark found

Tonic Textual's Graph Synthesis pipeline led all seven pipelines on every metric. It pairs Textual's NER with graph-based synthesis that links every mention of a person or organization across the export and replaces them consistently. The six other pipelines are pure-LLM agents: each model wrote its own extraction and rendering code in a sandbox, then performed NER and synthesis itself.

PrivacyBench v2 results: macro average across 21 datasets, ordered by the combined score

Privacy-bench results by pipeline: NER recall, synthesis accuracy, and combined synthesis + NER accuracy
Pipeline NER recall Synthesis accuracy Synthesis + NER accuracy
Textual Graph Synthesis 93.9% 91.8% 86.2%
GPT-6 Astra 91.6% 87.5% 80.1%
Claude Opus 5 93.6% 85.5% 80.0%
GPT-5.6 Sol 89.9% 81.6% 73.4%
Claude Sonnet 5 89.0% 79.3% 70.6%
GPT-5.6 Terra 84.9% 78.6% 66.7%
GPT-5.6 Luna 54.7% 77.4% 42.3%

Two results explain Textual's lead.

Textual's NER model is a small, efficient encoder-based model, not an LLM. That makes it much faster and cheaper to run than asking an LLM to find the PII, and on PrivacyBench it still detects PII better than the frontier models do. End to end, it caught slightly more of the known PII than any pure-LLM pipeline, and against the human annotations it posted the highest F1 of any engine tested, as the detection results below show.

The even more definitive lead is in synthesis. Finding a person's name is one job. Giving that person the same new identity in every email, Slack message, and document they appear in is a harder one, and it's where the pure-LLM pipelines lose ground. Textual's graph-based synthesis links every mention of a person or organization across the export before it replaces anything, so its replacements stayed coherent more often than any other pipeline's: 91.8% of the time, against 87.5% for the next best.

Put those two together and you get the clearest result in the benchmark. Textual Graph Synthesis used GPT-5.6 Sol for its LLM steps, which makes the GPT-5.6 Sol row a direct comparison: the same model, with and without Textual. Running the whole job on its own, Sol scored 73.4% and finished third among the pure-LLM pipelines, behind GPT-6 Astra and Claude Opus 5. Inside Textual, the same model scored 86.2%, a 12.8-point gain that took it from third place to first. Nothing about the model changed; what changed was the system around it.

Combined synthesis + NER accuracy across the seven PrivacyBench v2 pipelines. Textual Graph Synthesis leads at 86.2%.
Fig. 1: Combined synthesis + NER accuracy across the seven pipelines, macro average over 21 datasets.

How Textual's detection compares on human-annotated data

The end-to-end results measure recall against the generated answer key, which covers the seed characters and their organizations but not every name that appears in the data. To score detection completely, human annotators labeled every PII span in the emails, Slack messages, PDFs, and Word documents of 6 of the 21 datasets: 25,918 spans across 5,198 messages and 656 document pages. Against that human gold, precision and F1 become meaningful.

NER scores against the human annotations (overlap matching)

Privacy-bench NER detection quality by engine: precision, recall, and F1
NER engine Precision Recall F1
Textual via SDK 90.5% 93.0% 91.7%
Textual Graph Synthesis 88.0% 94.3% 91.0%
Claude Opus 5 92.2% 84.8% 88.3%
Claude Sonnet 5 92.3% 82.3% 87.0%
GPT-6 Astra 88.7% 85.0% 86.8%
GPT-5.6 Sol 90.3% 82.3% 86.1%
GPT-5.6 Terra 92.5% 78.2% 84.8%
Presidio 79.1% 87.0% 82.8%
GLiNER2 83.9% 68.0% 75.1%

Tonic Textual posted the highest F1 of any engine, 91.7%, and it's the only engine above 90% on precision, recall, and F1 at once. Some LLMs scored up to 2 points higher on precision, but they gave up more than 8 points of recall to get there. For privacy, recall is the number that matters: a precision miss over-redacts something that wasn't PII, while a recall miss leaves real PII in your data. The two open-source engines trailed further, with Presidio at 82.8% F1 and GLiNER2 at 75.1%.

The Textual NER models were not trained on any PrivacyBench data, and the Textual via SDK row is reproducible with a Textual API key and the open scoring code.

NER F1 against the human annotations. The two Textual engines lead every LLM and open-source engine.
Fig. 2: NER F1 (overlap matching) against the human annotations, by engine.

PrivacyBench v1 scored email and Slack messages supplied as pre-extracted text, across five entity types, with two-stage pipelines that ran NER and then synthesis. v2 keeps the v1 email and Slack messages, now in their raw export formats, along with the same three metric definitions, and turns the task into a full de-identification job:

  • Raw exports instead of extracted text, so pipelines do their own text extraction and rendering.
  • 357 documents in PDF, Word, Excel, and CSV formats, with ground truth located in each file's native coordinates.
  • Ten entity types instead of five, adding phone numbers, location addresses, employee IDs, account numbers, and URLs.
  • Human-annotated ground truth for six datasets, which makes full precision, recall, and F1 scoring possible for detection.
  • New pipelines: Textual Graph Synthesis and six agent pipelines built on current frontier models.

Because the task is harder, v2 scores run lower than v1 scores, and the two versions aren't comparable. Comparing them is like comparing a 5K time to a trail-race time: both measure how fast you are, on very different ground. The v1 data and scoring code remain available at the v1 revision of the dataset and the v1 tag of the metrics repo. We've also re-run v1 on the current version of Textual, so the v1 results below reflect how Textual performs today.

We've re-run PrivacyBench v1 using Textual Graph Synthesis, which is used for PrivacyBench v2. At the time of publishing v1, Textual Graph Synthesis was still in development.

On v1, Textual Graph Synthesis reached 95.4% on the combined synthesis + NER score, with 98.2% NER recall. The original v1 pipelines appear below as first published, alongside the new run.

PrivacyBench v1 results: the current Textual run and the original v1 pipelines

Privacy-bench results by pipeline and run: NER recall, synthesis accuracy, and combined synthesis + NER accuracy
Pipeline Run NER recall Synthesis accuracy Synthesis + NER accuracy
Textual Graph Synthesis Textual Graph Synthesis, September 2026 98.2% 97.2% 95.4%
Textual NER + Opus 4.8 synthesis Original v1, July 2026 95.0% 97.0% 92.0%
Opus 4.8 only Original v1, July 2026 88.8% 99.0% 87.7%
Textual NER + Sonnet 4.6 synthesis Original v1, July 2026 95.0% 90.5% 85.9%
Textual NER + Haiku 4.5 synthesis Original v1, July 2026 95.0% 87.3% 82.8%
Sonnet 4.6 only Original v1, July 2026 83.8% 90.5% 76.2%
Haiku 4.5 only Original v1, July 2026 79.9% 88.5% 71.1%

The Opus-only pipeline still posts the highest raw synthesis accuracy on v1, 99.0%, but only because its lower recall left it replacing the easier entities it caught. With detection misses counted, its combined score is 87.7%.

PrivacyBench v1: combined synthesis + NER accuracy for the current Textual run and the original v1 pipelines.
PrivacyBench v1: combined accuracy for the current Textual run and the original v1 pipelines.

Detection on v1's human-annotated messages

Human annotators fully labeled the email and Slack messages in 6 of the 21 datasets. That makes it possible to score v1 detection on precision, recall, and F1, rather than recall alone. These are the same 6 datasets used in v2. 

NER scores against the v1 human annotations for overlap metric

Privacy-bench NER detection quality by engine: precision, recall, and F1
NER engine Precision Recall F1
Textual Graph Synthesis 95.8% 96.3% 96.0%
Textual via SDK 96.5% 92.6% 94.5%
Opus 4.8 (LLM NER) 95.4% 88.2% 91.7%
Sonnet 4.6 (LLM NER) 97.0% 85.0% 90.6%
Presidio 87.3% 89.0% 88.1%
Haiku 4.5 (LLM NER) 97.6% 79.9% 87.8%
GLiNER2 85.2% 88.6% 86.9%

Textual is the clear winner for overall NER when precision, recall, and F1 are all considered across the six datasets. For both scoring methods, Textual's precision, recall, and F1 are all over 90%. It is the only model that achieves over 90% for all three scores, and it's the only model that achieves over 90% recall, which it does for both scoring methods. There are other models that score similarly to Textual on precision, but Textual wins by a lot on recall in those cases. Recall is the more important metric when considering privacy as it measures privacy leaks.

Put your own de-identification pipeline to the test

If you're de-identifying workspace data for AI, PrivacyBench v2 points one way: frontier LLMs working alone trail on keeping identities coherent across a full export, and Tonic Textual's graph-based synthesis closes that gap. The combined score, the same-model gain, and the detection results against human annotations all point the same direction.

The strongest case is the one you run yourself. Point your pipeline at the 21 public workspace exports, score it with the open metrics code, and compare your numbers with the ones here. To reproduce the Textual NER results, create a free Textual account for an API key and follow the instructions on the dataset card.

Tonic Textual is also the engine behind Tonic.ai's company data de-identification work, which prepares operational data like email, chat, tickets, and documents for licensing and model training. For the principles behind that work, read the core tenets of de-identification for AI data licensing. To see Textual on your own data, book a demo.

Frequently asked questions

v2 turns PrivacyBench into a full de-identification job. Pipelines now receive raw export files instead of extracted text, handle PDF, Word, Excel, and CSV documents alongside email and Slack messages, detect ten entity types instead of five, and write the synthetic export back out in its original formats. v2 also adds human-annotated ground truth for six datasets.

v2 turns PrivacyBench into a full de-identification job. Pipelines now receive raw export files instead of extracted text, handle PDF, Word, Excel, and CSV documents alongside email and Slack messages, detect ten entity types instead of five, and write the synthetic export back out in its original formats. v2 also adds human-annotated ground truth for six datasets.

No. v2 is a harder task with different data, so scores drop across the board and the two versions aren't comparable. Compare pipelines within a version, not across versions.

Yes. Against the human annotations, Tonic Textual posted the highest F1 of any engine tested, 91.7%, ahead of every frontier LLM and both open-source engines. Some LLMs scored slightly higher on precision but more than 8 points lower on recall, and recall is what determines whether real PII stays in your data.

Synthesis accuracy measures replacement quality: of the PII a pipeline detected and replaced, how much is coherent? An LLM judge checks each replacement, per character, against the synthetic identity it should belong to, so a nickname has to map to a nickname form of the same synthetic name, and a synthetic email address has to belong to the synthetic person.

Yes. The 21 workspace exports and their answer keys are public on Hugging Face, and the scoring code is on GitHub. Run your pipeline on the raw exports, write its replacements in the benchmark's output format, and score them against the same ground truth and metrics used here.

A public PII benchmark can't contain real PII, and trustworthy labels are expensive to annotate by hand. Because Tonic Fabricate generates each dataset from a seed that fixes every character's PII in advance, the ground truth comes with the data, and the difficulty can be dialed up on purpose. Human annotators then validated that answer key on six of the datasets.

Joe Ferrara is Head of AI at Tonic.ai, where he uses the latest developments in artificial intelligence to improve named entity recognition and synthetic data generation in Tonic Textual. He holds a Ph.D. in Mathematics from the University of California, Santa Cruz, and a B.A. in Mathematics from the University of California, Berkeley. Prior to his role at Tonic.ai, Joe served as a Data Scientist at ICW Group, where he gained experience applying traditional data science techniques in the context of the insurance industry. His background in theoretical mathematics research gives him a unique perspective on his work in artificial intelligence and machine learning.