Decide AI model ranking: intelligence, benchmarks and price AI model ranking: intelligence, benchmarks and price Each dot is a model: the higher, the better it ranks; the further left, the cheaper it is. Click a model to see it alone.
All models Open weights API only
Most attractive area 140 150 160 170 €0.10 €0.30 €1.00 €3.00 €10.00 €30.00 Price per million tokens (€, 3/4 input + 1/4 output, log scale) Epoch Capabilities Index (ECI) Claude Opus 5.5 (Anthropic) · Score 167.3 · €3.56 / €17.82 Claude Opus 5.5 Claude Sonnet 5.5 (Anthropic) · Score 165.2 · €1.78 / €8.91 Claude Sonnet 5.5 Claude Haiku 4.5 (Anthropic) · Score 142.4 · €0.89 / €4.45 Claude Haiku 4.5 GPT-5.6 Sol (OpenAI) · Score 161.8 · €3.56 / €17.82 GPT-5.6 Sol Gemini 3.8 Flash (Google) · Score 156.9 · €0.67 / €3.34 Gemini 3.8 Flash Grok 4.5 (xAI) · Score 154.0 · €1.78 / €5.35 Grok 4.5 GLM-5.3 (Z.ai) · Score 155.8 · €1.25 / €3.92 GLM-5.3 GLM-5.3 Flash (Z.ai) · Score 151.9 · €0.13 / €0.45 GLM-5.3 Flash DeepSeek V4.1 Flash (DeepSeek) · Score 155.0 · €0.27 / €1.07 DeepSeek V4.1 Flash DeepSeek V4 Pro (DeepSeek) · Score 155.4 · €1.18 / €3.53 DeepSeek V4 Pro Mistral Medium 3.5 (Mistral AI) · Score 141.4 · €1.34 / €6.68 Mistral Medium 3.5 GPT-6 Astra (OpenAI) · Score 166.5 · €8.91 / €44.54 GPT-6 Astra Claude Fable 5.1 (Anthropic) · Score 164.8 · €8.91 / €44.54 Claude Fable 5.1 Gemini 3.7 Flash (Google) · Score 157.4 · €0.67 / €3.34 Gemini 3.7 Flash Gemini 3.5 Flash-Lite (Google) · Score 145.1 · €0.27 / €2.23 Gemini 3.5 Flash-Lite Grok 4.6 (xAI) · Score 156.6 · €1.78 / €5.35 Grok 4.6 Open weights API only Pareto line Claude Opus 5.5 🇺🇸 Anthropic · Open weights: no · License: Proprietary
See all models Score 167.3 #1 of 37 on score · confidence interval 164.0–172.0
Price per million tokens · blended €7.13 #13 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 90.6 %
SWE-bench Verified — not measured
FrontierMath 91.2 %
SimpleQA Verified 72.2 %
MATH Level 5 — not measured
On the Pareto line: no model is both better and cheaper.
Claude Sonnet 5.5 🇺🇸 Anthropic · Open weights: no · License: Proprietary
See all models Score 165.2 #3 of 37 on score · confidence interval 161.8–169.4
Price per million tokens · blended €3.56 #12 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 95.6 %
SWE-bench Verified — not measured
FrontierMath 88.8 %
SimpleQA Verified 46.5 %
MATH Level 5 — not measured
On the Pareto line: no model is both better and cheaper.
Claude Haiku 4.5 🇺🇸 Anthropic · Open weights: no · License: Proprietary
See all models Score 142.4 #30 of 37 on score · confidence interval 139.5–144.3
Price per million tokens · blended €1.78 #7 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 71.2 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified 13.2 %
MATH Level 5 96.4 %
Off the Pareto line: beaten on both score and price by 6.
GPT-5.6 Sol 🇺🇸 OpenAI · Open weights: no · License: Proprietary
See all models Score 161.8 #5 of 37 on score · confidence interval 159.3–165.3
Price per million tokens · blended €7.13 #13 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 93.5 %
SWE-bench Verified — not measured
FrontierMath 89.1 %
SimpleQA Verified 69.7 %
MATH Level 5 — not measured
promotional price, at least until 11/21/2026
Off the Pareto line: beaten on both score and price by 1.
GPT-5.6 Terra 🇺🇸 OpenAI · Open weights: no · License: Proprietary
See all models Score 159.8 #6 of 37 on score · confidence interval 157.1–162.8
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 93.3 %
SWE-bench Verified — not measured
FrontierMath 86.0 %
SimpleQA Verified 43.2 %
MATH Level 5 — not measured
GPT-5.6 Luna 🇺🇸 OpenAI · Open weights: no · License: Proprietary
See all models Score 156.4 #13 of 37 on score · confidence interval 154.1–158.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 91.6 %
SWE-bench Verified — not measured
FrontierMath 82.1 %
SimpleQA Verified 41.0 %
MATH Level 5 — not measured
Gemini 3.8 Flash 🇺🇸 Google · Open weights: no · License: Proprietary
See all models Score 156.9 #9 of 37 on score · confidence interval 154.6–160.4
Price per million tokens · blended €1.34 #4 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 95.4 %
SWE-bench Verified — not measured
FrontierMath 68.4 %
SimpleQA Verified 69.7 %
MATH Level 5 — not measured
until 12/31/2026, then $1.50 / $7.50
Off the Pareto line: beaten on both score and price by 0.
Muse Spark 1.3 🇺🇸 Meta · Open weights: no · License: Proprietary
See all models Score 156.9 #9 of 37 on score · confidence interval 154.7–159.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond — not measured
SWE-bench Verified — not measured
FrontierMath 74.4 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
Grok 4.5 🇺🇸 xAI · Open weights: no · License: Proprietary
See all models Score 154.0 #17 of 37 on score · confidence interval 152.3–156.1
Price per million tokens · blended €2.67 #10 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 93.4 %
SWE-bench Verified — not measured
FrontierMath 57.2 %
SimpleQA Verified 48.3 %
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 5.
Qwen3.8 Max 🇨🇳 Alibaba · Open weights: no · License: Proprietary
See all models Score 156.6 #11 of 37 on score · confidence interval 154.5–158.8
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 92.7 %
SWE-bench Verified — not measured
FrontierMath 74.7 %
SimpleQA Verified 45.8 %
MATH Level 5 — not measured
GLM-5.3 🇨🇳 Z.ai · Open weights: yes · License: MIT
See all models Score 155.8 #14 of 37 on score · confidence interval 153.7–158.3
Price per million tokens · blended €1.92 #8 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 90.9 %
SWE-bench Verified — not measured
FrontierMath 68.8 %
SimpleQA Verified 41.0 %
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 2.
Resources to run it yourself →
GLM-5.3 Flash 🇨🇳 Z.ai · Open weights: yes · License: MIT
See all models Score 151.9 #18 of 37 on score · confidence interval 149.4–154.3
Price per million tokens · blended €0.21 #1 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 90.2 %
SWE-bench Verified — not measured
FrontierMath 55.8 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
On the Pareto line: no model is both better and cheaper.
Resources to run it yourself →
DeepSeek V4.1 Flash 🇨🇳 DeepSeek · Open weights: yes · License: MIT
See all models Score 155.0 #16 of 37 on score · confidence interval 148.8–157.6
Price per million tokens · blended €0.47 #2 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond — not measured
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
peak-hour price, half price off-peak
On the Pareto line: no model is both better and cheaper.
Resources to run it yourself →
DeepSeek V4 Pro 🇨🇳 DeepSeek · Open weights: yes · License: MIT
See all models Score 155.4 #15 of 37 on score · confidence interval 153.7–157.6
Price per million tokens · blended €1.76 #6 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 91.7 %
SWE-bench Verified — not measured
FrontierMath 64.6 %
SimpleQA Verified 52.9 %
MATH Level 5 — not measured
peak-hour price, half price off-peak
Off the Pareto line: beaten on both score and price by 2.
Resources to run it yourself →
Qwen3.8 27B 🇨🇳 Alibaba · Open weights: yes · License: Apache 2.0
See all models Score 149.4 #21 of 37 on score · confidence interval 147.4–151.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond — not measured
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Qwen3.5 397B 🇨🇳 Alibaba · Open weights: yes · License: Apache 2.0
See all models Score 146.7 #24 of 37 on score · confidence interval 144.8–148.2
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 86.4 %
SWE-bench Verified — not measured
FrontierMath 31.2 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
MiniMax M3 🇨🇳 MiniMax · Open weights: yes · License: MiniMax Community License
See all models Score 147.0 #23 of 37 on score · confidence interval 142.7–149.8
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 90.9 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Gemma 4 31B 🇺🇸 Google · Open weights: yes · License: Apache 2.0
See all models Score 142.8 #29 of 37 on score · confidence interval 140.3–144.9
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 75.8 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified 10.4 %
MATH Level 5 — not measured
Resources to run it yourself →
Mistral Medium 3.5 🇫🇷 Mistral AI · Open weights: yes · License: Modified MIT
See all models Score 141.4 #32 of 37 on score · confidence interval 138.4–143.7
Price per million tokens · blended €2.67 #9 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond — not measured
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 8.
Resources to run it yourself →
Mistral Small 3.2 🇫🇷 Mistral AI · Open weights: yes · License: Apache 2.0
See all models Score 131.7 #37 of 37 on score · confidence interval 126.6–133.9
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 49.1 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
gpt-oss-120b 🇺🇸 OpenAI · Open weights: yes · License: Apache 2.0
See all models Score 139.9 #33 of 37 on score · confidence interval 135.3–142.3
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 75.8 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
gpt-oss-20b 🇺🇸 OpenAI · Open weights: yes · License: Apache 2.0
See all models Score 137.8 #35 of 37 on score · confidence interval 133.0–139.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 60.8 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Llama 4 Maverick 🇺🇸 Meta · Open weights: yes · License: Llama 4
See all models Score 132.2 #36 of 37 on score · confidence interval 128.0–134.1
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 67.0 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 73.0 %
Resources to run it yourself →
GPT-6 Astra 🇺🇸 OpenAI · Open weights: no · License: Proprietary
See all models Score 166.5 #2 of 37 on score · confidence interval 163.1–171.1
Price per million tokens · blended €17.82 #15 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 95.8 %
SWE-bench Verified — not measured
FrontierMath 93.7 %
SimpleQA Verified 75.6 %
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 1.
Claude Fable 5.1 🇺🇸 Anthropic · Open weights: no · License: Proprietary
See all models Score 164.8 #4 of 37 on score · confidence interval 161.8–169.1
Price per million tokens · blended €17.82 #15 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond — not measured
SWE-bench Verified — not measured
FrontierMath 90.2 %
SimpleQA Verified 70.8 %
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 2.
Gemini 3.7 Flash 🇺🇸 Google · Open weights: no · License: Proprietary
See all models Score 157.4 #8 of 37 on score · confidence interval 155.4–160.1
Price per million tokens · blended €1.34 #4 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 94.8 %
SWE-bench Verified — not measured
FrontierMath 71.6 %
SimpleQA Verified 69.2 %
MATH Level 5 — not measured
until 12/31/2026, then $1.50 / $7.50
On the Pareto line: no model is both better and cheaper.
Gemini 3.5 Flash-Lite 🇺🇸 Google · Open weights: no · License: Proprietary
See all models Score 145.1 #26 of 37 on score · confidence interval 142.5–146.7
Price per million tokens · blended €0.76 #3 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 83.3 %
SWE-bench Verified — not measured
FrontierMath 26.0 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 2.
Grok 4.6 🇺🇸 xAI · Open weights: no · License: Proprietary
See all models Score 156.6 #11 of 37 on score · confidence interval 154.7–158.9
Price per million tokens · blended €2.67 #10 of 16 on price
Benchmarks (best result measured by Epoch) GPQA Diamond 94.0 %
SWE-bench Verified — not measured
FrontierMath 66.0 %
SimpleQA Verified 49.3 %
MATH Level 5 — not measured
Off the Pareto line: beaten on both score and price by 2.
Qwen3.7 Flash 🇨🇳 Alibaba · Open weights: no · License: Proprietary
See all models Score 144.6 #27 of 37 on score · confidence interval 142.4–147.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 82.3 %
SWE-bench Verified — not measured
FrontierMath 19.3 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
Kimi K3 🇨🇳 Moonshot AI · Open weights: yes · License: Kimi K3 license
See all models Score 157.6 #7 of 37 on score · confidence interval 154.9–160.4
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 93.1 %
SWE-bench Verified — not measured
FrontierMath 72.2 %
SimpleQA Verified 50.6 %
MATH Level 5 — not measured
Resources to run it yourself →
Kimi K2.7 Code 🇨🇳 Moonshot AI · Open weights: yes · License: —
See all models Score 150.0 #20 of 37 on score · confidence interval 148.1–151.8
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 87.9 %
SWE-bench Verified — not measured
FrontierMath 54.0 %
SimpleQA Verified 36.5 %
MATH Level 5 — not measured
Resources to run it yourself →
Inkling 🇺🇸 Thinking Machines · Open weights: yes · License: Apache 2.0
See all models Score 148.6 #22 of 37 on score · confidence interval 145.7–150.6
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 88.3 %
SWE-bench Verified — not measured
FrontierMath 33.3 %
SimpleQA Verified 40.3 %
MATH Level 5 — not measured
Resources to run it yourself →
Inkling Small 🇺🇸 Thinking Machines · Open weights: yes · License: Apache 2.0
See all models Score 150.2 #19 of 37 on score · confidence interval 147.5–152.1
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 88.5 %
SWE-bench Verified — not measured
FrontierMath 46.3 %
SimpleQA Verified 19.1 %
MATH Level 5 — not measured
Resources to run it yourself →
Nemotron 3 Ultra 🇺🇸 Nvidia · Open weights: yes · License: OpenMDW-1.1
See all models Score 146.2 #25 of 37 on score · confidence interval 143.9–148.1
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 85.4 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Qwen3.6 35B-A3B 🇨🇳 Alibaba · Open weights: yes · License: —
See all models Score 143.9 #28 of 37 on score · confidence interval 141.2–146.0
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 84.8 %
SWE-bench Verified — not measured
FrontierMath 20.4 %
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Gemma 4 26B A4B 🇺🇸 Google · Open weights: yes · License: Apache 2.0
See all models Score 141.8 #31 of 37 on score · confidence interval 138.4–143.4
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 73.2 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Qwen3.5 9B 🇨🇳 Alibaba · Open weights: yes · License: —
See all models Score 139.5 #34 of 37 on score · confidence interval 136.4–141.3
Price per million tokens — No official price published by the developer: this model is not on the chart.
Benchmarks (best result measured by Epoch) GPQA Diamond 79.0 %
SWE-bench Verified — not measured
FrontierMath — not measured
SimpleQA Verified — not measured
MATH Level 5 — not measured
Resources to run it yourself →
Model Developer Open weights License Score Input / output price Blended price GPQA Diamond Claude Opus 5.5 Pareto 🇺🇸 Anthropic no Proprietary 167.3 €3.56 / €17.82 €7.13 90.6 % GPT-6 Astra 🇺🇸 OpenAI no Proprietary 166.5 €8.91 / €44.54 €17.82 95.8 % Claude Sonnet 5.5 Pareto 🇺🇸 Anthropic no Proprietary 165.2 €1.78 / €8.91 €3.56 95.6 % Claude Fable 5.1 🇺🇸 Anthropic no Proprietary 164.8 €8.91 / €44.54 €17.82 — GPT-5.6 Sol 🇺🇸 OpenAI no Proprietary 161.8 €3.56 / €17.82 €7.13 93.5 % GPT-5.6 Terra 🇺🇸 OpenAI no Proprietary 159.8 not published — 93.3 % Kimi K3 🇨🇳 Moonshot AI yes Kimi K3 license 157.6 not published — 93.1 % Gemini 3.7 Flash Pareto 🇺🇸 Google no Proprietary 157.4 €0.67 / €3.34 €1.34 94.8 % Gemini 3.8 Flash 🇺🇸 Google no Proprietary 156.9 €0.67 / €3.34 €1.34 95.4 % Muse Spark 1.3 🇺🇸 Meta no Proprietary 156.9 not published — — Qwen3.8 Max 🇨🇳 Alibaba no Proprietary 156.6 not published — 92.7 % Grok 4.6 🇺🇸 xAI no Proprietary 156.6 €1.78 / €5.35 €2.67 94.0 % GPT-5.6 Luna 🇺🇸 OpenAI no Proprietary 156.4 not published — 91.6 % GLM-5.3 🇨🇳 Z.ai yes MIT 155.8 €1.25 / €3.92 €1.92 90.9 % DeepSeek V4 Pro 🇨🇳 DeepSeek yes MIT 155.4 €1.18 / €3.53 €1.76 91.7 % DeepSeek V4.1 Flash Pareto 🇨🇳 DeepSeek yes MIT 155.0 €0.27 / €1.07 €0.47 — Grok 4.5 🇺🇸 xAI no Proprietary 154.0 €1.78 / €5.35 €2.67 93.4 % GLM-5.3 Flash Pareto 🇨🇳 Z.ai yes MIT 151.9 €0.13 / €0.45 €0.21 90.2 % Inkling Small 🇺🇸 Thinking Machines yes Apache 2.0 150.2 not published — 88.5 % Kimi K2.7 Code 🇨🇳 Moonshot AI yes — 150.0 not published — 87.9 % Qwen3.8 27B 🇨🇳 Alibaba yes Apache 2.0 149.4 not published — — Inkling 🇺🇸 Thinking Machines yes Apache 2.0 148.6 not published — 88.3 % MiniMax M3 🇨🇳 MiniMax yes MiniMax Community License 147.0 not published — 90.9 % Qwen3.5 397B 🇨🇳 Alibaba yes Apache 2.0 146.7 not published — 86.4 % Nemotron 3 Ultra 🇺🇸 Nvidia yes OpenMDW-1.1 146.2 not published — 85.4 % Gemini 3.5 Flash-Lite 🇺🇸 Google no Proprietary 145.1 €0.27 / €2.23 €0.76 83.3 % Qwen3.7 Flash 🇨🇳 Alibaba no Proprietary 144.6 not published — 82.3 % Qwen3.6 35B-A3B 🇨🇳 Alibaba yes — 143.9 not published — 84.8 % Gemma 4 31B 🇺🇸 Google yes Apache 2.0 142.8 not published — 75.8 % Claude Haiku 4.5 🇺🇸 Anthropic no Proprietary 142.4 €0.89 / €4.45 €1.78 71.2 % Gemma 4 26B A4B 🇺🇸 Google yes Apache 2.0 141.8 not published — 73.2 % Mistral Medium 3.5 🇫🇷 Mistral AI yes Modified MIT 141.4 €1.34 / €6.68 €2.67 — gpt-oss-120b 🇺🇸 OpenAI yes Apache 2.0 139.9 not published — 75.8 % Qwen3.5 9B 🇨🇳 Alibaba yes — 139.5 not published — 79.0 % gpt-oss-20b 🇺🇸 OpenAI yes Apache 2.0 137.8 not published — 60.8 % Llama 4 Maverick 🇺🇸 Meta yes Llama 4 132.2 not published — 67.0 % Mistral Small 3.2 🇫🇷 Mistral AI yes Apache 2.0 131.7 not published — 49.1 %
FAQ What does “open weights” mean? Weights are the billions of numbers that hold what a model has learned. An open-weight model can be downloaded and run on your own machine; a closed model is only available through its developer's API. Open weights does not mean open source: the license may restrict use.
How is model intelligence measured? With Epoch AI's Capabilities Index (ECI), which combines the results of many benchmarks: scientific reasoning (GPQA Diamond), coding (SWE-bench Verified), mathematics (FrontierMath), factual accuracy (SimpleQA Verified).
Where do the prices come from? From the developers' official pricing pages, in dollars per million tokens, converted to euros at the European Central Bank reference rate. The blended price counts 3/4 of the input price and 1/4 of the output price.
What is the Pareto line? It links the models that no other model beats on both intelligence and price: the best trade-offs right now.
Was this information useful?