BenchLeader
CohereOpen weightsAuto-detected

Command A 03

Best configuration ranks #436 of 610 on the BenchLeader Index at 43.5 ±6.8. Last measured 2 Sept 2026. Released 13 Mar 2025.

Blended price
$4.38/M
$2.50 in · $10.00 out
Output speed
First answer
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index44
  2. Reasoning53
  3. Coding31
  4. Maths35
  5. Knowledge38
  6. Human preference53

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
LMArena Hard Prompts1368#176LMArena
GPQA Diamond (Vals)not in index48.5%#118Vals AI

Coding

BenchmarkdefaultSource
LMArena Coding1389#182LMArena
LiveCodeBench35.1%#130Vals AI
SWE-bench (Vals)not in index7.8%#84Vals AI

Maths

BenchmarkdefaultSource
AIME (Vals)13.3%#84Vals AI
MGSM85.7%#64Vals AI

Knowledge

BenchmarkdefaultSource
MMLU-Pro69.2%#117Vals AI
LegalBench79.7%#85Vals AI
CorpFin46.0%#108Vals AI
TaxEval61.4%#118Vals AI
MedQA80.5%#69Vals AI

Human preference

BenchmarkdefaultSource
LMArena Text1354#175LMArena

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0040
Summarise a 30-page report12,000 / 600$0.036
Code edit6,000 / 1,500$0.030
Agentic coding session60,000 / 4,000$0.190
Structured extraction2,000 / 200$0.0070

See also

Data as of 9 Sept 2026. Compare with another model.