MMLU, GPQA, SWE-bench, HumanEval, AIME, MATH, LiveBench, SimpleBench, and Arena Elo supported.
Benchmark Center
AI Benchmark Center
Compare public AI benchmark scores with AI Pricing Hub pricing data, cost per benchmark point, and sourced model rankings.
Coverage
Sourced benchmark coverage
Rows with explicit public source URLs.
Models joined to pricing data.
Missing benchmark values are left blank.
Best value
Transparent Value Score leaders
| Rank | Model | Provider | Avg score | Cost | Benchmark / dollar | Scores |
|---|---|---|---|---|---|---|
| 1 | gpt-4.1-nano | OpenAI | 65.2 | $0.5 | 130.40 | MMLU 80.1 GPQA 50.3 |
| 2 | gpt-4o | OpenAI | 77.3 | $12.5 | 6.18 | MMLU 88.7 GPQA 53.6 MATH 76.6 HumanEval 90.2 |
| 3 | claude-3-opus-20240229 | Anthropic | 68.6 | $90 | 0.76 | MMLU 86.8 GPQA 50.4 |
Explorer
Price / Performance Explorer
Rankings
Best value rankings
Overall
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
Coding
- #1gpt-4o82.72/100
Reasoning
- #1claude-3-opus-2024022921.98/100
Vision
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
RAG
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
Translation
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
Long Context
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
Enterprise
- #1gpt-4o82.72/100
- #2gpt-4.1-nano63.75/100
- #3claude-3-opus-2024022921.98/100
Visualizations
Performance vs cost
Benchmark pages
Task-specific benchmark rankings
Best coding models
Rank sourced benchmark scores against token pricing.
Best reasoning models
Rank sourced benchmark scores against token pricing.
Best math models
Rank sourced benchmark scores against token pricing.
Best vision models
Rank sourced benchmark scores against token pricing.
Best multilingual models
Rank sourced benchmark scores against token pricing.
Best value models
Rank sourced benchmark scores against token pricing.
Providers
Provider benchmark pages
Anthropic
Compare sourced benchmark rows for this provider with current pricing.
Cohere
Compare sourced benchmark rows for this provider with current pricing.
DeepSeek
Compare sourced benchmark rows for this provider with current pricing.
Google Gemini
Compare sourced benchmark rows for this provider with current pricing.
Groq
Compare sourced benchmark rows for this provider with current pricing.
OpenAI
Compare sourced benchmark rows for this provider with current pricing.
OpenRouter
Compare sourced benchmark rows for this provider with current pricing.
xAI
Compare sourced benchmark rows for this provider with current pricing.
Historical intelligence
History-derived market signals
These insights are generated only from stored pricing history snapshots across 534 tracked model histories.
349 tracked models
100.0/100 average stability
466.8% cumulative largest drops
77 deprecated
100/100
85/100
$0.04 latest combined price
32 tracked days
FAQ
Benchmark FAQ
Does AI Pricing Hub create benchmark scores?
No. Benchmark rows are only shown when they are present in the sourced benchmark dataset.
What does cost per benchmark point mean?
It divides the model combined input plus output price per 1M tokens by the sourced benchmark score.
Why are some benchmark cells blank?
Blank cells mean no sourced public score has been added for that benchmark and model combination.
Are benchmark scores directly comparable?
Not always. Benchmark methodology, prompting, dates, and provider reporting can differ. Use source links before making decisions.
Continue
Pick up where you left off
Public API
Build with AI Pricing Hub data
Use static JSON endpoints for providers, models, rankings, history, market metrics, and changelog events.
Editorial information
Reviewed by AI Pricing Hub Editorial
2026-08-04
Methodology explains collection, validation, limitations, and update cadence.