AI evaluation of English text anonymization
Checks that an anonymization model or context lets no U.S. personal data through (email, phone, SSN, date of birth, medical record number).
Evaluated resource
Criteria
Each piece of personal data must be replaced by the right tag; a text without personal data must stay identical.
Test cases
| # | Input | Expected answer |
|---|---|---|
| 1 | Call Daniel Brooks at 312-555-0199. | \[PHONE_1\] |
| 2 | My email is k.nguyen@outlook.com, thanks. | \[EMAIL_1\] |
| 3 | SSN: 078-05-1120 | \[SSN_1\] |
| 4 | Born on 07/22/1985, Sarah Mitchell lives in Denver. | \[DOB_1\] |
| 5 | MRN 55120034 was admitted on Monday. | \[MRN_1\] |
| 6 | The package leaves Seattle tomorrow. | ^The package leaves Seattle tomorrow\.$ |
Estimated cost
Input tokens per call : ≈ 320
| Model | Per call | Per 1,000 calls |
|---|---|---|
| Mistral Large · Mistral AI | €0.00014 | €0.14 |
| DeepSeek Flash · DeepSeek | €0.00009 | €0.08 |
| DeepSeek V4 Pro · DeepSeek | €0.00037 | €0.37 |
| GPT-6 Luna · OpenAI | €0.00003 | €0.03 |
| GPT-6 Sol · OpenAI | €0.00056 | €0.56 |
| Claude Sonnet 5.5 · Anthropic | €0.00056 | €0.56 |
| Claude Opus 5.5 · Anthropic | €0.00113 | €1.13 |
| Self-hosted open-source model (Ollama, vLLM) | €0 in API fees (server cost only) | |
Input cost only, excluding the model's answer. Providers' public standard prices (Mistral AI, DeepSeek, OpenAI, Anthropic) converted from USD to euros at the ECB rate of September 29, 2026. Estimate: 1 token ≈ 3.6 characters.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data: no summary required.
United States · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data: no documentation to publish.
China · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistics
Likes : 0 · Comments : 0
Click a counter to show or hide it, hover the chart for details. One view per visitor per day, bots excluded; counting started on 30 September 2026.
- Identifier
- contextetech--eval-anonymize-english-text
- Type
- Evaluations
- File
- eval-anonymize-english-text.eval.jsonl
- Format
- JSON Lines (application/jsonl)
- Size
- 476 bytes
- Encoding
- UTF-8
- Estimated tokens
- ≈ 132
- Language
- English
- License
- CC BY 4.0
- Commercial use
- Allowed
- Personal data
- Fictional (made up)
- Version
- 1.0
- Published on
- October 2, 2026 at 9:54 PM
- Updated on
- October 2, 2026 at 9:54 PM
- Storage
- in database
- Uses
- 0
- Likes
- 0
eval-anonymize-english-text.eval.jsonl
{"input":"Call Daniel Brooks at 312-555-0199.","expected":"\\[PHONE_1\\]"}
{"input":"My email is k.nguyen@outlook.com, thanks.","expected":"\\[EMAIL_1\\]"}
{"input":"SSN: 078-05-1120","expected":"\\[SSN_1\\]"}
{"input":"Born on 07/22/1985, Sarah Mitchell lives in Denver.","expected":"\\[DOB_1\\]"}
{"input":"MRN 55120034 was admitted on Monday.","expected":"\\[MRN_1\\]"}
{"input":"The package leaves Seattle tomorrow.","expected":"^The package leaves Seattle tomorrow\\.$"}
Run the evaluation (Python)
import json, re
cases = [json.loads(l) for l in open("eval-anonymize-english-text.eval.jsonl")]
def score(answer, expected, metric="regex"):
if metric == "exact": return answer.strip() == expected.strip()
if metric == "regex": return re.search(expected, answer) is not None
return expected.lower() in answer.lower() # contains (llm-judge : faire noter par un modèle)
def run(ask): # ask(input) -> réponse du modèle
ok = sum(score(ask(c["input"]), c["expected"]) for c in cases)
print(f"{ok}/{len(cases)} réussis ({100 * ok // len(cases)} %), seuil 100 %")
Community
No comments yet. Share your feedback, it will help the next person.