BenchLeader

Gemini 3 Flash vs Muse Spark 1.3

Verdict
  • Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.6.
  • Gemini 3 Flash (thinking) is stronger in agents & tools, instruction following.
  • Muse Spark 1.3 (xhigh) is stronger in composite, knowledge, long context, multimodal, coding.
  • Gemini 3 Flash (thinking) is 1.8× cheaper ($1.13 vs $2.00 per 1M blended).
MetricGemini 3 Flash (thinking)Muse Spark 1.3 (xhigh)
BenchLeader Index61.663.9
Agents & tools score67.460.2
Composite score62.378.6
Instruction following score75.4
Knowledge score69.075.1
Long context score65.568.1
Multimodal score64.466.6
Coding score70.3
Blended price $/M$1.13$2.00
Output speed180 tok/s190 tok/s
Time to first answer6.1 s35.7 s
Context window1M1M
LMArena WebDev1625
LiveBench81.6%
LiveBench Reasoning89.7%
LiveBench Coding81.1%
LiveBench Agentic Coding64.1%
LiveBench Mathematics96.0%
LiveBench Data Analysis79.6%
LiveBench Language82.8%
AA Intelligence Index26.345.2
IFBench78.0%
AA-LCR78.0%83.0%
MMMU-Pro79.9%82.0%
AA-Omniscience10.123.1
Terminal-Bench Hard38.6%
GPQA Diamond (AA)89.8%94.1%
Humanity's Last Exam (AA)36.6%47.5%
SciCode (AA)59.7%
τ²-Bench Telecom (AA)80.4%
Terminal-Bench 2.1 (Vals)72.3%
Vals Index60.3

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Gemini 3 Flash vs Muse Spark 1.3: questions

Is Gemini 3 Flash better than Muse Spark 1.3?
Muse Spark 1.3 (xhigh) leads on quality: 63.9 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Muse Spark 1.3 (xhigh) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Gemini 3 Flash better than Muse Spark 1.3 for agentic tasks?
Gemini 3 Flash scores higher in agentic tasks (67 vs 60 on the category index, where 50 is average).
Which is cheaper, Gemini 3 Flash or Muse Spark 1.3?
Gemini 3 Flash is cheaper: $1.13 against $2.00 per million tokens, blended at three input tokens per output token.
Which is faster, Gemini 3 Flash or Muse Spark 1.3?
Muse Spark 1.3 streams faster: 190 against 180 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.