AI evaluation of English text anonymization
Checks that an anonymization model or context lets no U.S. personal data through (email, phone, SSN, date of birth, medical record number).
Recurso evaluado
Criterios
Each piece of personal data must be replaced by the right tag; a text without personal data must stay identical.
Casos de prueba
| # | Entrada | Respuesta esperada |
|---|---|---|
| 1 | Call Daniel Brooks at 312-555-0199. | \[PHONE_1\] |
| 2 | My email is k.nguyen@outlook.com, thanks. | \[EMAIL_1\] |
| 3 | SSN: 078-05-1120 | \[SSN_1\] |
| 4 | Born on 07/22/1985, Sarah Mitchell lives in Denver. | \[DOB_1\] |
| 5 | MRN 55120034 was admitted on Monday. | \[MRN_1\] |
| 6 | The package leaves Seattle tomorrow. | ^The package leaves Seattle tomorrow\.$ |
Coste estimado
Tokens de entrada por llamada : ≈ 320
| Modelo | Por llamada | Por 1.000 llamadas |
|---|---|---|
| Mistral Large · Mistral AI | 0,00014 € | 0,14 € |
| DeepSeek Flash · DeepSeek | 0,00009 € | 0,08 € |
| DeepSeek V4 Pro · DeepSeek | 0,00037 € | 0,37 € |
| GPT-6 Luna · OpenAI | 0,00003 € | 0,03 € |
| GPT-6 Sol · OpenAI | 0,00056 € | 0,56 € |
| Claude Sonnet 5.5 · Anthropic | 0,00056 € | 0,56 € |
| Claude Opus 5.5 · Anthropic | 0,00113 € | 1,13 € |
| Modelo de código abierto autoalojado (Ollama, vLLM) | 0 € de API (solo el coste del servidor) | |
Solo coste de entrada, sin la respuesta del modelo. Precios públicos estándar de los proveedores (Mistral AI, DeepSeek, OpenAI, Anthropic) convertidos de USD a euros al tipo del BCE del 29 de septiembre de 2026. Estimación: 1 token ≈ 3,6 caracteres.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data: no summary required.
United States · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data: no documentation to publish.
China · 5/5
Fictional: invented names, addresses and numbers, no real person.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Estadísticas
Me gusta : 0 · Comentarios : 0
Haz clic en un contador para mostrarlo u ocultarlo; pasa el ratón por el gráfico para ver el detalle. Una visita por visitante y día, sin robots; recuento iniciado el 30 de septiembre de 2026.
- Identificador
- contextetech--eval-anonymize-english-text
- Tipo
- Evaluaciones
- Archivo
- eval-anonymize-english-text.eval.jsonl
- Formato
- JSON Lines (application/jsonl)
- Tamaño
- 476 bytes
- Codificación
- UTF-8
- Tokens estimados
- ≈ 132
- Idioma
- Inglés
- Licencia
- CC BY 4.0
- Uso comercial
- Permitido
- Datos personales
- Ficticios (inventados)
- Versión
- 1.0
- Publicado el
- 2 de octubre de 2026 a las 21:54
- Actualizado el
- 2 de octubre de 2026 a las 21:54
- Almacenamiento
- en base de datos
- Usos
- 0
- Me gusta
- 0
eval-anonymize-english-text.eval.jsonl
{"input":"Call Daniel Brooks at 312-555-0199.","expected":"\\[PHONE_1\\]"}
{"input":"My email is k.nguyen@outlook.com, thanks.","expected":"\\[EMAIL_1\\]"}
{"input":"SSN: 078-05-1120","expected":"\\[SSN_1\\]"}
{"input":"Born on 07/22/1985, Sarah Mitchell lives in Denver.","expected":"\\[DOB_1\\]"}
{"input":"MRN 55120034 was admitted on Monday.","expected":"\\[MRN_1\\]"}
{"input":"The package leaves Seattle tomorrow.","expected":"^The package leaves Seattle tomorrow\\.$"}
Ejecutar la evaluación (Python)
import json, re
cases = [json.loads(l) for l in open("eval-anonymize-english-text.eval.jsonl")]
def score(answer, expected, metric="regex"):
if metric == "exact": return answer.strip() == expected.strip()
if metric == "regex": return re.search(expected, answer) is not None
return expected.lower() in answer.lower() # contains (llm-judge : faire noter par un modèle)
def run(ask): # ask(input) -> réponse du modèle
ok = sum(score(ask(c["input"]), c["expected"]) for c in cases)
print(f"{ok}/{len(cases)} réussis ({100 * ok // len(cases)} %), seuil 100 %")
Comunidad
Aún no hay comentarios. Comparte tu opinión, ayudará a los siguientes.