BenchLeader

gpt-oss-120b

Best configuration ranks #283 of 610 on the BenchLeader Index at 50.5 ±7.7 (high reasoning effort). Last measured 5 Aug 2025stale: no new result in six months. Released 5 Aug 2025.

Blended price
Output speed
202 tok/s
First answer
11 s
first token 0.49 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index51
  2. Reasoning56
  3. Coding46
  4. Maths51

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low46.0#376224 tok/s9.68 s$0.0002Agents & tools 38 · Coding 39 · Composite 42 · Instruction following 58 · Knowledge 39 · Long context 49
medium202 tok/s11 sCoding 45
highbest50.5#283202 tok/s11 sCoding 46 · Maths 51 · Reasoning 56
default47.6#348202 tok/s11 s$0.0002Agents & tools 37 · Coding 43 · Composite 45 · Human preference 53 · Instruction following 51 · Knowledge 46 · Long context 52 · Maths 60 · Reasoning 41

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSource
GPQA Diamond75.8%#131Epoch AI Benchmarking Hub
SimpleBench22.1%#78SimpleBench
LMArena Hard Prompts1362#181LMArena
GPQA Diamond (AA)not in index67.2%#30778.2%#207Artificial Analysis
Humanity's Last Exam (AA)not in index5.9%#33819.6%#162Artificial Analysis
GPQA Diamond (Vals)not in index78.5%#70Vals AI
Kagi LLM Benchmark58.6%#53Kagi LLM Benchmark

Coding

BenchmarklowmediumhighdefaultSource
SciCode36.0%#12738.9%#116SciCode
WeirdML41.9%#9348.2%#70WeirdML
LMArena Coding1391#180LMArena
SciCode (AA)not in index34.0%#138Artificial Analysis
LiveCodeBench83.2%#49Vals AI
SWE-bench (Vals)not in index33.6%#83Vals AI
SWE-Bench Pro16.2%#19Scale AI SEAL
Aider Polyglot41.8%#28Aider polyglot leaderboard
SWE-bench Verified (bash only)26.0%#38SWE-bench
SWE-bench Verified (any scaffold)not in index26.0%#57SWE-bench

Agents & tools

BenchmarklowmediumhighdefaultSource
Terminal-Bench18.7%#57Terminal-Bench
APEX-Agents4.7%#57Mercor
Terminal-Bench Hard5.3%#26323.5%#143Artificial Analysis
τ²-Bench Telecom (AA)not in index45.0%#20665.8%#161Artificial Analysis

Maths

BenchmarklowmediumhighdefaultSource
OTIS Mock AIME88.9%#61Epoch AI Benchmarking Hub
AIME (Vals)92.6%#18Vals AI
MGSM92.0%#24Vals AI
MathArena Apex1.0%#35MathArena

Knowledge

BenchmarklowmediumhighdefaultSource
AA-Omniscience-53.5#347-49.3#311Artificial Analysis
MMLU-Pro79.2%#94Vals AI
LegalBench75.9%#103Vals AI
CorpFin58.2%#82Vals AI
TaxEval71.6%#74Vals AI
MedQA91.4%#38Vals AI
PRBench Finance43.8%#18Scale AI SEAL
PRBench Legal40.2%#24Scale AI SEAL
MultiNRC15.2%#38Scale AI SEAL

Instruction following

BenchmarklowmediumhighdefaultSource
IFBench58.3%#12969.0%#75Artificial Analysis
MultiChallenge45.3%#24Scale AI SEAL

Human preference

BenchmarklowmediumhighdefaultSource
LMArena Text1352#178LMArena

Long context

BenchmarklowmediumhighdefaultSource
AA-LCR46.0%#26452.0%#250Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSource
Epoch Capabilities Indexnot in index140.1#99Epoch AI Benchmarking Hub
AA Intelligence Index10.2#30912.3#268Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Cerebras696 tok/s0.19 s$0.350$0.750131kfp16
Amazon Bedrock293 tok/s0.41 s$0.150$0.600131k
SambaNova256 tok/s0.68 s$0.140$0.950131k
Groq250 tok/s0.27 s$0.150$0.600131k
DeepInfra (fp8)166 tok/s1.94 s$0.200$0.950131kfp8
Baseten (US)157 tok/s0.30 s$0.100$0.500128kfp4
Baseten155 tok/s0.31 s$0.100$0.500128kfp4
Google Vertex140 tok/s0.34 s$0.090$0.360131k
Nebius Token Factory140 tok/s0.54 s$0.150$0.600131kfp4
Amazon Bedrock (EU)132 tok/s0.43 s$0.150$0.600131k
DeepInfra (Turbo)121 tok/s0.42 s$0.150$0.600131kbf16
MARA94 tok/s3.02 s$0.150$0.750131k
Parasail91 tok/s0.44 s$0.100$0.750131kfp4
Phala71 tok/s1.19 s$0.150$0.600131k
Together64 tok/s0.30 s$0.150$0.600131k
NovitaAI61 tok/s1.17 s$0.050$0.250131kfp4
AkashML47 tok/s1.32 s$0.030$0.170131kbf16
Mancer36 tok/s0.61 s$0.055$0.500131kfp8
CoreWeave33 tok/s0.45 s$0.030$0.170131kfp4
DigitalOcean23 tok/s0.75 s$0.055$0.385128k
DeepInfra (bf16)22 tok/s2.67 s$0.037$0.170131kbf16
SiliconFlow10 tok/s5.42 s$0.050$0.450131kfp8

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.151$0.303$0.454$0.605Jun 26Jul 26Aug 26
input outputnow $0.050 in · $0.500 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.075 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 30012.2 s
Summarise a 30-page report12,000 / 60013.7 s
Code edit6,000 / 1,50018.1 s
Agentic coding session60,000 / 4,00030.5 s
Structured extraction2,000 / 20011.7 s

See also

Data as of 9 Sept 2026. Compare these configurations.