BenchLeader

Gemini 4 Argon vs Muse Spark 1.2

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 62.7.
  • Gemini 4 Argon (high) is stronger in agents & tools, coding, composite, human preference, knowledge, long context, maths, reasoning.
  • Muse Spark 1.2 (xhigh) is stronger in multimodal.
  • Muse Spark 1.2 (xhigh) is 2.0× cheaper ($2.00 vs $4.00 per 1M blended).
MetricGemini 4 Argon (high)Muse Spark 1.2 (xhigh)
BenchLeader Index70.062.7
Agents & tools score67.555.9
Coding score72.155.0
Composite score89.266.6
Human preference score73.269.1
Knowledge score75.766.9
Long context score65.064.7
Maths score71.966.7
Reasoning score82.069.8
Multimodal score–65.3
Blended price $/M$4.00$2.00
Output speed–325 tok/s
Time to first answer–18.2 s
Context window1M1.0M
SimpleQA Verified–60.3%
SciCode61.8%56.4%
WeirdML–60.3%
LMArena Text15251492
LMArena Hard Prompts15511508
LMArena Coding15621531
LMArena WebDev16781533
LMArena Vision–1305
LMArena Agent9.3-3.7
LiveBench–78.0%
LiveBench Reasoning–90.0%
LiveBench Coding–77.5%
LiveBench Agentic Coding–57.6%
LiveBench Mathematics–91.2%
LiveBench Data Analysis–76.5%
LiveBench Language–78.6%
LiveBench Instruction Following–74.3%
AA Intelligence Index v4.3.252.639.6
AA-LCR79.7%79.0%
AA-Omniscience42.427.2
GPQA Diamond (AA)–90.4%
Humanity's Last Exam (AA)57.1%45.5%
SciCode (AA)61.8%57.4%
MMLU-Pro–88.3%
IOI100.0%21.8%
LegalBench88.3%85.3%
CorpFin–70.9%
TaxEval–80.4%
Terminal-Bench 2.1 (Vals)–69.7%
SWE-bench (Vals)–86.6%
Vals Index68.949.3
CritPt27.1%17.7%
GDPval-AA v2.156.3%48.9%
τ³-Banking (AA)–34.9%
BioMysteryBench76.3%64.8%
Code Migration68.2%29.9%
CUA-bench4.8%–
CyberBench77.9%69.5%
Excel Modeling Benchmark75.2%57.0%
Finance Agent v265.4%60.6%
Harvey's Legal Agent Benchmark19.6%25.4%
Legal Research Bench54.8%43.8%
MedCode58.8%49.4%
MedScribe87.4%90.1%
MMMU-Pro (Vals)–86.1%
MortgageTax–65.4%
MysteryMechanism45.5%–
ProgramBench2.5%–
Public Benefits Bench69.8%68.5%
SAGE53.6%47.7%
SkillsBench–53.0%
SREBench44.3%–
Tax Agent Bench76.2%56.9%
Terminal-Bench 4.0 (Vals)57.6%6.1%
Terminal-Bench Science44.3%–
Vibe Code Bench v1.191.9%79.1%
LMArena Maths15281481
LMArena Creative Writing15191454
LMArena Instruction Following15291471
LMArena Multi-turn15531506
LMArena Longer Queries15451489
DeepSWE v1.1–54.9%
LMCA–48.4%
DTBench–94.7%
GDP.pdf–16.0%
FrontierSWE55.0%12.0%
Terminal-Bench 4.0 (AA)57.1%7.1%
Terminal-Bench 2.1 (AA)–80.2%
AutomationBench77.5%–
GDP.pdf21.8%–
AA-Omniscience: accuracy49.9%45.4%
AA-Omniscience: non-hallucination84.9%66.7%
AA-Briefcase v1.11488–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 4 Argon vs Muse Spark 1.2: questions

Is Gemini 4 Argon better than Muse Spark 1.2?
Gemini 4 Argon (high) leads on quality: 70.0 vs 62.7. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 4 Argon better than Muse Spark 1.2 for coding?
Gemini 4 Argon scores higher in coding (72 vs 55 on the category index, where 50 is average).
Is Gemini 4 Argon better than Muse Spark 1.2 for agentic tasks?
Gemini 4 Argon scores higher in agentic tasks (68 vs 56 on the category index, where 50 is average).
Which is cheaper, Gemini 4 Argon or Muse Spark 1.2?
Muse Spark 1.2 is cheaper: $2.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Muse Spark 1.2 accepts more context: 1.0M against 1M tokens.