BenchLeader
MetaAuto-detected

Muse Spark

Best configuration ranks #27 of 610 on the BenchLeader Index at 65.4 ±3.3. Last measured 2 Sept 2026. Released 8 Apr 2026.

Blended price
Output speed
First answer
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index65
  2. Reasoning70
  3. Coding66
  4. Agents & tools66
  5. Maths56
  6. Knowledge65
  7. Instruction following80
  8. Human preference69
  9. Multimodal66
  10. Long context66
  11. Composite68

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking
defaultbest65.4#27Agents & tools 66 · Coding 66 · Composite 68 · Human preference 69 · Instruction following 80 · Knowledge 65 · Long context 66 · Maths 56 · Multimodal 66 · Reasoning 70

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkthinkingdefaultSource
SciCode51.5%#52SciCode
LMArena Coding1526#18LMArena
SWE-bench (Vals)not in index74.4%#46Vals AI
SWE-Bench Pro55.0%#3Scale AI SEAL

Agents & tools

BenchmarkthinkingdefaultSource
Terminal-Bench Hard45.5%#32Artificial Analysis
τ²-Bench Telecom (AA)not in index91.5%#58Artificial Analysis
MCP Atlas82.2%#7Scale AI SEAL

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME88.9%#60Epoch AI Benchmarking Hub
ProofBench17.0%#40Vals AI
AIME (Vals)96.9%#2Vals AI

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience7.2#67Artificial Analysis
MMLU-Pro87.3%#30Vals AI
LegalBench84.2%#34Vals AI
CorpFin65.1%#33Vals AI
TaxEval77.7%#3Vals AI
PRBench Finance52.4%#4Scale AI SEAL
PRBench Legal52.3%#3Scale AI SEAL
MultiNRC59.0%#5Scale AI SEAL

Instruction following

BenchmarkthinkingdefaultSource
IFBench75.9%#24Artificial Analysis
MultiChallenge75.5%#1Scale AI SEAL
TutorBench68.5%#1Scale AI SEAL

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1488#13LMArena

Multimodal

BenchmarkthinkingdefaultSource
LMArena Vision1306#9LMArena
MMMU-Pro80.5%#31Artificial Analysis
MMMU-Pro (official)not in index80.4%#3MMMU

Long context

BenchmarkthinkingdefaultSource
AA-LCR78.0%#82Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
Epoch Capabilities Indexnot in index152.1#35Epoch AI Benchmarking Hub
AA Intelligence Index31.3#73Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

See also

Data as of 9 Sept 2026. Compare these configurations.