AI Pricing Hub Benchmark Center

Benchmark Center

AI Benchmark Center

Compare public AI benchmark scores with AI Pricing Hub pricing data, cost per benchmark point, and sourced model rankings.

Coverage

Sourced benchmark coverage

Benchmark types9

MMLU, GPQA, SWE-bench, HumanEval, AIME, MATH, LiveBench, SimpleBench, and Arena Elo supported.

Sourced records8

Rows with explicit public source URLs.

Models with scores3

Models joined to pricing data.

No fabricated scoresEnforced

Missing benchmark values are left blank.

Best value

Transparent Value Score leaders

RankModelProviderAvg scoreCostBenchmark / dollarScores
1gpt-4.1-nanoOpenAI65.2$0.5130.40MMLU 80.1
GPQA 50.3
2gpt-4oOpenAI77.3$12.56.18MMLU 88.7
GPQA 53.6
MATH 76.6
HumanEval 90.2
3claude-3-opus-20240229Anthropic68.6$900.76MMLU 86.8
GPQA 50.4

Explorer

Price / Performance Explorer

Price vs performance
Price / performance scatter plotX axis is combined input plus output price per 1M tokens. Y axis is average sourced benchmark performance. Bubble size is context window. Color is provider.claude-3-opus-20240229 | price $90 | performance 68.6 | value score 21.98gpt-4.1-nano | price $0.5 | performance 65.2 | value score 63.75gpt-4o | price $12.5 | performance 77.3 | value score 82.72PricePerformance
AnthropicOpenAI

Rankings

Best value rankings

Best value

Coding

  1. #1gpt-4o82.72/100

Visualizations

Performance vs cost

Performance vs cost
Performance vs costgpt-4.1-nano 65.2 at $0.5gpt-4o 77.3 at $12.5claude-3-opus-20240229 68.6 at $90CostScore
Cost per benchmark ranking
Cost per benchmark rankinggpt-4.1-nano65.2gpt-4o77.3claude-3-opus-2024022968.6

Benchmark pages

Task-specific benchmark rankings

Benchmark ranking

Best math models

Rank sourced benchmark scores against token pricing.

Benchmark ranking

Best value models

Rank sourced benchmark scores against token pricing.

Providers

Provider benchmark pages

Provider benchmarks

Anthropic

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

Cohere

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

DeepSeek

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

Google Gemini

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

Groq

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

OpenAI

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

OpenRouter

Compare sourced benchmark rows for this provider with current pricing.

Provider benchmarks

xAI

Compare sourced benchmark rows for this provider with current pricing.

Historical intelligence

History-derived market signals

These insights are generated only from stored pricing history snapshots across 534 tracked model histories.

Fastest growing providerOpenRouter

349 tracked models

Most stable providerCohere

100.0/100 average stability

Aggressive reductionsOpenRouter

466.8% cumulative largest drops

Most deprecated modelsOpenRouter

77 deprecated

Most stable model~anthropic/claude-fable-latest

100/100

Most volatile modeldeepseek/deepseek-v4-flash-0731

85/100

Cheapest over timeinclusionai/ling-2.6-flash

$0.04 latest combined price

Longest supportedgpt-5.4-mini

32 tracked days

Price volatility by provider
Price volatility by providerGenerated only from AI Pricing Hub history snapshots.OpenRouterxAI4.05
Historical provider activity
Historical provider activityGenerated only from AI Pricing Hub history snapshots.AnthropicxAI493.00

Open stability rankings

FAQ

Benchmark FAQ

Does AI Pricing Hub create benchmark scores?

No. Benchmark rows are only shown when they are present in the sourced benchmark dataset.

What does cost per benchmark point mean?

It divides the model combined input plus output price per 1M tokens by the sourced benchmark score.

Why are some benchmark cells blank?

Blank cells mean no sourced public score has been added for that benchmark and model combination.

Are benchmark scores directly comparable?

Not always. Benchmark methodology, prompting, dates, and provider reporting can differ. Use source links before making decisions.

Public API

Build with AI Pricing Hub data

Use static JSON endpoints for providers, models, rankings, history, market metrics, and changelog events.

Newsletter

Get AI pricing changes in your inbox

Monthly pricing moves, new model launches, and practical cost notes. Provider integration is not enabled yet.

Editorial information

Reviewed by AI Pricing Hub Editorial

Last updated

2026-08-04

Methodology

Methodology explains collection, validation, limitations, and update cadence.