BenchLeader

Gemini 4 Argon vs Muse Spark 1.3

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 67.7.
  • Gemini 4 Argon (high) is stronger in coding, composite, human preference, knowledge, maths, reasoning.
  • Muse Spark 1.3 (max) is stronger in agents & tools, long context, multimodal.
  • Muse Spark 1.3 (max) is 2.0× cheaper ($2.00 vs $4.00 per 1M blended).
MetricGemini 4 Argon (high)Muse Spark 1.3 (max)
BenchLeader Index70.067.7
Agents & tools score67.570.3
Coding score72.164.6
Composite score89.284.1
Human preference score73.269.4
Knowledge score75.772.9
Long context score65.066.8
Maths score71.963.6
Reasoning score82.078.1
Multimodal score–65.9
Blended price $/M$4.00$2.00
Output speed–175 tok/s
Time to first answer–70.5 s
Context window1M1.0M
FrontierMath Tiers 1–3–74.0%
FrontierMath Tier 4–46.3%
SciCode61.8%58.8%
ProofBench–58.0%
LMArena Text15251494
LMArena Hard Prompts15511517
LMArena Coding15621538
LMArena WebDev16781657
LMArena Vision–1309
LMArena Agent9.34
AA Intelligence Index v4.3.252.648.1
AA-LCR79.7%83.0%
AA-Omniscience42.425
GPQA Diamond (AA)–93.5%
Humanity's Last Exam (AA)57.1%48.7%
SciCode (AA)61.8%58.8%
IOI100.0%56.6%
LegalBench88.3%–
Terminal-Bench 2.1 (Vals)–79.0%
Vals Index68.958.2
CritPt27.1%24.9%
GDPval-AA v2.156.3%59.0%
τ³-Banking (AA)–50.5%
ITBench SRE (AA)–33.2%
BioMysteryBench76.3%–
Code Migration68.2%47.4%
CUA-bench4.8%5.8%
CyberBench77.9%72.7%
Excel Modeling Benchmark75.2%67.4%
Finance Agent v265.4%60.0%
Harvey's Legal Agent Benchmark19.6%23.8%
Legal Research Bench54.8%55.3%
MedCode58.8%–
MedScribe87.4%–
MysteryMechanism45.5%36.0%
ProgramBench2.5%–
Public Benefits Bench69.8%–
SAGE53.6%–
SREBench44.3%–
Tax Agent Bench76.2%72.4%
Terminal-Bench 4.0 (Vals)57.6%24.8%
Terminal-Bench Science44.3%10.0%
Vibe Code Bench 1-100–20.5%
Vibe Code Bench v1.191.9%85.9%
LMArena Maths15281502
LMArena Creative Writing15191456
LMArena Instruction Following15291483
LMArena Multi-turn15531487
LMArena Longer Queries15451497
LMArena Document–1468
Chess Puzzles–38.0%
Mystery Game Puzzles–25.0%
CursorBench–41.6%
FrontierSWE55.0%–
Terminal-Bench 4.0 (AA)57.1%33.3%
Terminal-Bench 2.1 (AA)–84.3%
AutomationBench77.5%57.9%
GDP.pdf21.8%26.6%
MLCR–43.3%
Harvey LAB–8.9%
AA-Omniscience: accuracy49.9%43.6%
AA-Omniscience: non-hallucination84.9%67.1%
AA-Briefcase v1.114881581

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Gemini 4 Argon vs Muse Spark 1.3: questions

Is Gemini 4 Argon better than Muse Spark 1.3?
Gemini 4 Argon (high) leads on quality: 70.0 vs 67.7. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Gemini 4 Argon better than Muse Spark 1.3 for coding?
Gemini 4 Argon scores higher in coding (72 vs 65 on the category index, where 50 is average).
Is Gemini 4 Argon better than Muse Spark 1.3 for agentic tasks?
Muse Spark 1.3 scores higher in agentic tasks (70 vs 68 on the category index, where 50 is average).
Which is cheaper, Gemini 4 Argon or Muse Spark 1.3?
Muse Spark 1.3 is cheaper: $2.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Muse Spark 1.3 accepts more context: 1.0M against 1M tokens.