BenchLeader

gpt-oss-20b

Best configuration ranks #355 of 610 on the BenchLeader Index at 47.1 ±4.6. Last measured 2 Sept 2026. Released 5 Aug 2025.

Blended price
$0.085/M
$0.050 in · $0.190 out
Output speed
192 tok/s
First answer
11 s
first token 0.43 s
Context
131k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index47
  2. Reasoning49
  3. Coding55
  4. Agents & tools29
  5. Maths55
  6. Knowledge37
  7. Instruction following64
  8. Human preference48
  9. Long context43
  10. Composite40

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low44.6#411249 tok/s8.87 s$0.0001Agents & tools 38 · Composite 42 · Instruction following 58 · Knowledge 36 · Long context 41 · Maths 42 · Reasoning 40
medium192 tok/s11 sCoding 41 · Maths 53 · Reasoning 45
high192 tok/s11 sCoding 40 · Maths 46 · Reasoning 36
defaultbest47.1#355192 tok/s11 s$0.0001Agents & tools 29 · Coding 55 · Composite 40 · Human preference 48 · Instruction following 64 · Knowledge 37 · Long context 43 · Maths 55 · Reasoning 49

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSource
GPQA Diamond53.2%#19460.8%#17746.0%#217Epoch AI Benchmarking Hub
LMArena Hard Prompts1322#219LMArena
GPQA Diamond (AA)not in index61.1%#35068.8%#296Artificial Analysis
Humanity's Last Exam (AA)not in index5.3%#36111.0%#240Artificial Analysis
GPQA Diamond (Vals)not in index68.9%#94Vals AI
Kagi LLM Benchmark53.2%#71Kagi LLM Benchmark

Coding

BenchmarklowmediumhighdefaultSource
SciCode34.4%#136SciCode
WeirdML36.8%#11840.9%#99WeirdML
LMArena Coding1369#200LMArena
SciCode (AA)not in index38.9%#121Artificial Analysis
LiveCodeBench80.4%#67Vals AI

Agents & tools

BenchmarklowmediumhighdefaultSource
Terminal-Bench3.4%#66Terminal-Bench
Terminal-Bench Hard4.5%#27010.6%#213Artificial Analysis
τ²-Bench Telecom (AA)not in index50.3%#19060.2%#174Artificial Analysis

Maths

BenchmarklowmediumhighdefaultSource
OTIS Mock AIME40.3%#17965.3%#13850.8%#166Epoch AI Benchmarking Hub
AIME (Vals)86.0%#33Vals AI
MGSM89.0%#48Vals AI

Knowledge

BenchmarklowmediumhighdefaultSource
AA-Omniscience-58.5#378-63.0#410Artificial Analysis
MMLU-Pro71.6%#111Vals AI
LegalBench70.8%#112Vals AI
CorpFin53.1%#95Vals AI
TaxEval63.7%#115Vals AI
MedQA82.9%#63Vals AI
MultiNRC10.4%#40Scale AI SEAL

Instruction following

BenchmarklowmediumhighdefaultSource
IFBench57.8%#13365.1%#103Artificial Analysis

Human preference

BenchmarklowmediumhighdefaultSource
LMArena Text1317#218LMArena

Long context

BenchmarklowmediumhighdefaultSource
AA-LCR31.0%#32234.7%#305Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSource
Epoch Capabilities Indexnot in index137.8#107Epoch AI Benchmarking Hub
AA Intelligence Index9.9#3139.0#338Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Groq384 tok/s0.54 s$0.075$0.300131k
Amazon Bedrock381 tok/s0.32 s$0.070$0.150131k
Google Vertex259 tok/s0.38 s$0.070$0.250131k
CoreWeave104 tok/s0.11 s$0.030$0.130131kfp4
DeepInfra85 tok/s0.31 s$0.030$0.140131kbf16
NovitaAI77 tok/s0.78 s$0.040$0.150131kfp4
Parasail66 tok/s0.48 s$0.030$0.150131kfp4
Phala64 tok/s0.28 s$0.040$0.150131k
SiliconFlow55 tok/s1.07 s$0.040$0.180131kfp8
Together48 tok/s0.30 s$0.050$0.200131k
Darkbloom28 tok/s3.28 s$0.020$0.100131kfp8
AkashML28 tok/s1.56 s$0.020$0.100131kfp4

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.055$0.110$0.165$0.220Jan 26Feb 26Mar 26Apr 26May 26Jun 26Jul 26Aug 26
input outputnow $0.030 in · $0.150 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.037 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001$0.000112.8 s
Summarise a 30-page report12,000 / 600$0.0007$0.000614.3 s
Code edit6,000 / 1,500$0.0006$0.000519.0 s
Agentic coding session60,000 / 4,000$0.0038$0.003232.0 s
Structured extraction2,000 / 200$0.0001$0.000112.2 s

See also

Data as of 9 Sept 2026. Compare these configurations.