Which AI model for RAG over your documents?
An assistant that answers from your documents (RAG): product documentation, procedures, knowledge base.
DeepSeek V4.1 Flash
€0.27 / €1.07 per million tokens (input / output)
On the Pareto line, capability score of 155 at a low price: suited to the long passages read for each question.
Cost for 1 million questions of about 4,000 tokens read (question and passages) and 400 written: €1,496.
See in the ranking →Qwen3.8 27B
Free to use: only your hardware cost (a 24 GB GPU is enough with Q4 quantization, embeddings included).
Open weights under the Apache 2.0 license: with a local embedding model, neither documents nor questions leave your server.
See in the ranking →Claude Sonnet 5.5
€1.78 / €8.91 per million tokens (input / output)
On the Pareto line, with one of the best capability scores: better at saying “I don't know” when the passages don't hold the answer.
Cost for 1 million questions of about 4,000 tokens read (question and passages) and 400 written: €10,692.
See in the ranking →Why these choices
For RAG, quality depends first on retrieving the right passages (the embedding model and how documents are chunked), then on the model that writes the answer. Each question reads several thousand tokens: the input price matters more than the output price.
Take action
Data: Epoch AI index (CC BY 4.0), developers' official prices converted at the ECB rate, updated 2026-10-03 · AI model ranking: intelligence, benchmarks and price