Which AI model to extract data from documents?
Invoices, resumes, tickets, forms: turning text into structured data (JSON).
GLM-5.3 Flash
€0.13 / €0.45 per million tokens (input / output)
The cheapest model on the Pareto line: no model is both more capable and cheaper.
Cost for 1 million documents of about 1,500 tokens read and 300 written: €335.
See in the ranking →Qwen3.8 27B
Free to use: only your hardware cost (a 24 GB GPU is enough with Q4 quantization).
Open weights under the Apache 2.0 license, 27 billion parameters: it runs on a single server with a 24 GB GPU, and nothing leaves your premises.
See in the ranking →Claude Sonnet 5.5
€1.78 / €8.91 per million tokens (input / output)
On the Pareto line, with one of the best capability scores: fewer errors to fix on complex or poorly scanned documents.
Cost for 1 million documents of about 1,500 tokens read and 300 written: €5,346.
See in the ranking →Why these choices
For extraction, what matters: following a JSON format, inventing nothing, and the price per document at high volume. No benchmark in the index measures extraction directly: the capability index is a general guide, and price makes the difference at scale.
Take action
Data: Epoch AI index (CC BY 4.0), developers' official prices converted at the ECB rate, updated 2026-10-03 · AI model ranking: intelligence, benchmarks and price