BenchLeader

GPT-6 Astra vs MiMo-V2.6-Pro

Verdict
  • GPT-6 Astra (max) leads on quality: 71.4 vs 64.3.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, human preference, knowledge, maths, multimodal.
  • MiMo-V2.6-Pro is stronger in composite, long context, reasoning.
  • MiMo-V2.6-Pro is 36× cheaper ($0.548 vs $20.00 per 1M blended).
MetricGPT-6 Astra (max)MiMo-V2.6-Pro
BenchLeader Index71.464.3
Agents & tools score68.455.1
Coding score77.0
Composite score81.984.8
Human preference score68.2
Knowledge score81.466.8
Long context score66.169.1
Maths score76.9
Multimodal score66.8
Reasoning score77.888.8
Blended price $/M$20.00$0.548
Output speed53 tok/s54 tok/s
Time to first answer361.5 s39.8 s
Context window1.1M1.0M
GPQA Diamond95.8%
FrontierMath Tiers 1–393.7%
FrontierMath Tier 497.6%
OTIS Mock AIME100.0%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SciCode56.5%
WeirdML93.3%
FrontierCode53.3%
LMArena Text1480
LMArena Hard Prompts1497
LMArena Coding1543
LMArena WebDev1793
LMArena Vision1279
LMArena Agent11.5
LiveBench82.2%
LiveBench Reasoning92.7%
LiveBench Coding80.4%
LiveBench Agentic Coding57.3%
LiveBench Mathematics96.8%
LiveBench Data Analysis83.0%
LiveBench Language89.4%
LiveBench Instruction Following75.6%
AA Intelligence Index52.746.3
AA-LCR80.7%86.3%
MMMU-Pro86.9%
AA-Omniscience43.48.4
GPQA Diamond (AA)96.1%
Humanity's Last Exam (AA)54.7%49.4%
SciCode (AA)56.5%60.9%
IOI100.0%
Terminal-Bench 2.1 (Vals)87.3%67.8%
Vals Index66.659.7
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-398.6%
CritPt31.7%26.6%
GDPval (AA)52.1%58.7%
τ²-Bench Banking (AA)41.4%
Analyst Agent (AA)51.3%
BioMysteryBench79.3%
Code Migration67.7%43.0%
CUA-bench19.2%
CyberBench41.1%
Excel Modeling Benchmark71.7%62.9%
Finance Agent v253.5%58.3%
Harvey's Legal Agent Benchmark5.4%10.8%
Legal Research Bench39.4%47.1%
MedCode48.5%
MedScribe87.9%
MysteryMechanism53.1%
ProgramBench5.5%
SAGE46.4%
SREBench56.9%
Tax Agent Bench63.3%
Terminal-Bench 4.0 (Vals)57.1%
Terminal-Bench Science65.7%2.9%
Time Horizon Index: KSP90.5%
Vibe Code Bench 1-10027.6%
Vibe Code Bench v1.189.6%85.2%
DrugDiscoveryBench68.7%
LMArena Creative Writing1461
LMArena Instruction Following1461
LMArena Multi-turn1499
LMArena Longer Queries1481
LMArena Document1473
Chess Puzzles72.0%
EBR-bench76.2%
Mystery Game Puzzles84.0%
BALROG68.3%
DeepSWE73.2%
ALE-Bench2951.3
FrontierSWE65.5%
Terminal-Bench 4.0 (AA)59.1%34.9%
Terminal-Bench 2.1 (AA)88.4%
AutomationBench68.5%58.6%
GDP.pdf31.0%19.2%
MLCR35.0%
AA-Omniscience: accuracy62.6%34.9%
AA-Omniscience: non-hallucination48.7%59.4%
AA-Briefcase15691522

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

GPT-6 Astra vs MiMo-V2.6-Pro: questions

Is GPT-6 Astra better than MiMo-V2.6-Pro?
GPT-6 Astra (max) leads on quality: 71.4 vs 64.3. The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-09-23, but check the category scores for your use.
Is GPT-6 Astra better than MiMo-V2.6-Pro for agentic tasks?
GPT-6 Astra scores higher in agentic tasks (68 vs 55 on the category index, where 50 is average).
Which is cheaper, GPT-6 Astra or MiMo-V2.6-Pro?
MiMo-V2.6-Pro is cheaper: $0.548 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6 Astra or MiMo-V2.6-Pro?
MiMo-V2.6-Pro streams faster: 54 against 53 output tokens per second.
Which has the larger context window?
GPT-6 Astra accepts more context: 1.1M against 1.0M tokens.