BenchLeader

GPT-5.3-Codex vs MiMo-V2.6-Pro

Verdict
  • GPT-5.3-Codex (xhigh) and MiMo-V2.6-Pro are level on quality (64.6 vs 64.3).
  • GPT-5.3-Codex (xhigh) is stronger in agents & tools, coding, instruction following, knowledge, multimodal.
  • MiMo-V2.6-Pro is stronger in composite, long context, reasoning.
  • MiMo-V2.6-Pro is 8.8× cheaper ($0.548 vs $4.81 per 1M blended).
  • GPT-5.3-Codex (xhigh) streams 2.7× faster (146 vs 54 tokens per second).
MetricGPT-5.3-Codex (xhigh)MiMo-V2.6-Pro
BenchLeader Index64.664.3
Agents & tools score79.855.1
Coding score61.1
Composite score68.284.8
Instruction following score73.4
Knowledge score67.966.8
Long context score67.569.1
Multimodal score62.2
Reasoning score71.588.8
Blended price $/M$4.81$0.548
Output speed146 tok/s54 tok/s
Time to first answer52.3 s39.8 s
Context window400k1.0M
WeirdML77.9%
AA Intelligence Index32.546.3
IFBench75.4%
AA-LCR83.3%86.3%
MMMU-Pro78.5%
AA-Omniscience10.98.4
Terminal-Bench Hard53.0%
GPQA Diamond (AA)91.5%
Humanity's Last Exam (AA)42.5%49.4%
SciCode (AA)60.9%
τ²-Bench Telecom (AA)86.0%
LiveCodeBench87.3%
IOI53.8%
Terminal-Bench 2.1 (Vals)67.8%
SWE-bench (Vals)78.0%
Vals Index59.7
CritPt16.9%26.6%
GDPval (AA)58.7%
Code Migration43.0%
Excel Modeling Benchmark62.9%
Finance Agent v258.3%
Harvey's Legal Agent Benchmark10.8%
Legal Research Bench47.1%
Terminal-Bench Science2.9%
Vibe Code Bench v1.161.8%85.2%
ALE-Bench1655.2
Terminal-Bench 4.0 (AA)34.9%
AutomationBench58.6%
GDP.pdf19.2%
AA-Omniscience: accuracy52.9%34.9%
AA-Omniscience: non-hallucination10.8%59.4%
AA-Briefcase1522

Data as of 2026-09-23. Best configuration of each model; every score links to its source on the model pages.

GPT-5.3-Codex vs MiMo-V2.6-Pro: questions

Is GPT-5.3-Codex better than MiMo-V2.6-Pro?
GPT-5.3-Codex (xhigh) and MiMo-V2.6-Pro are level on quality (64.6 vs 64.3). The BenchLeader Index combines every independent quality benchmark; GPT-5.3-Codex (xhigh) is ahead overall as of 2026-09-23, but check the category scores for your use.
Is GPT-5.3-Codex better than MiMo-V2.6-Pro for agentic tasks?
GPT-5.3-Codex scores higher in agentic tasks (80 vs 55 on the category index, where 50 is average).
Which is cheaper, GPT-5.3-Codex or MiMo-V2.6-Pro?
MiMo-V2.6-Pro is cheaper: $0.548 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.3-Codex or MiMo-V2.6-Pro?
GPT-5.3-Codex streams faster: 146 against 54 output tokens per second.
Which has the larger context window?
MiMo-V2.6-Pro accepts more context: 1.0M against 400k tokens.