EN
Sign in Publish

Which AI model to extract data from documents?

Invoices, resumes, tickets, forms: turning text into structured data (JSON).

Your constraints
High volume

GLM-5.3 Flash

Z.ai · License: MIT · Capability index 151.9

€0.13 / €0.45 per million tokens (input / output)

The cheapest model on the Pareto line: no model is both more capable and cheaper.

Cost for 1 million documents of about 1,500 tokens read and 300 written: €335.

See in the ranking →
Sensitive data

Qwen3.8 27B

Alibaba · License: Apache 2.0 · Capability index 149.4

Free to use: only your hardware cost (a 24 GB GPU is enough with Q4 quantization).

Open weights under the Apache 2.0 license, 27 billion parameters: it runs on a single server with a 24 GB GPU, and nothing leaves your premises.

See in the ranking →
Maximum accuracy

Claude Sonnet 5.5

Anthropic · License: Proprietary · Capability index 165.2

€1.78 / €8.91 per million tokens (input / output)

On the Pareto line, with one of the best capability scores: fewer errors to fix on complex or poorly scanned documents.

Cost for 1 million documents of about 1,500 tokens read and 300 written: €5,346.

See in the ranking →

Why these choices

For extraction, what matters: following a JSON format, inventing nothing, and the price per document at high volume. No benchmark in the index measures extraction directly: the capability index is a general guide, and price makes the difference at scale.

Take action

Data: Epoch AI index (CC BY 4.0), developers' official prices converted at the ECB rate, updated 2026-10-03 · AI model ranking: intelligence, benchmarks and price

Was this information useful?