BenchLeader
MetaReasoning modelAuto-detected

Muse Spark 1.2

Best configuration ranks #52 of 610 on the BenchLeader Index at 62.6 ±6.5. Last measured 9 Aug 2026. Released 5 Aug 2026.

Blended price
$2.00/M
$1.25 in · $4.25 out
Output speed
219 tok/s
First answer
24 s
first token 3.92 s
Context
1.0M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index63
  2. Reasoning70
  3. Coding67
  4. Maths54
  5. Knowledge77
  6. Long context66
  7. Composite79

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
xhigh61.3#68219 tok/s24 s$0.0018Agents & tools 51 · Coding 63 · Composite 60 · Human preference 70 · Knowledge 65 · Multimodal 66 · Reasoning 69
defaultbest62.6#52219 tok/s24 s$0.0018Coding 67 · Composite 79 · Knowledge 77 · Long context 66 · Maths 54 · Reasoning 70

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkxhighdefaultSourceTrend
SimpleBench74.5%#9SimpleBench
LMArena Hard Prompts1513#12LMArena
LiveBench Reasoningnot in index90.0%#8LiveBench
GPQA Diamond (AA)not in index90.4%#62Artificial Analysis
Humanity's Last Exam (AA)not in index45.5%#28Artificial Analysis

Coding

BenchmarkxhighdefaultSourceTrend
SciCode56.4%#16SciCode
WeirdML60.3%#47WeirdML
LMArena Coding1533#10LMArena
LMArena WebDev1534#31LMArena
LiveBench Codingnot in index77.5%#26LiveBench
SciCode (AA)not in index57.4%#13Artificial Analysis
IOI49.5%#9Vals AI
SWE-bench (Vals)not in index86.6%#13Vals AI

Agents & tools

BenchmarkxhighdefaultSourceTrend
LMArena Agent-1.6#28LMArena
LiveBench Agentic Codingnot in index57.6%#16LiveBench
Terminal-Bench 2.1 (Vals)69.7%#21Vals AI

Maths

BenchmarkxhighdefaultSourceTrend
ProofBench43.0%#23Vals AI
LiveBench Mathematicsnot in index91.2%#19LiveBench

Knowledge

BenchmarkxhighdefaultSourceTrend
SimpleQA Verified60.3%#12Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index76.5%#23LiveBench
AA-Omniscience27.2#26Artificial Analysis
MMLU-Pro88.3%#20Vals AI
LegalBench85.3%#20Vals AI
CorpFin70.9%#5Vals AI
TaxEval80.4%#1Vals AI

Instruction following

BenchmarkxhighdefaultSourceTrend
LiveBench Languagenot in index78.6%#29LiveBench

Human preference

BenchmarkxhighdefaultSourceTrend
LMArena Text1499#5LMArena

Multimodal

BenchmarkxhighdefaultSourceTrend
LMArena Vision1304#12LMArena

Long context

BenchmarkxhighdefaultSourceTrend
AA-LCR79.0%#71Artificial Analysis

Composite

BenchmarkxhighdefaultSourceTrend
Epoch Capabilities Indexnot in index155.5#20Epoch AI Benchmarking Hub
LiveBench78.0%#13LiveBench
AA Intelligence Index39.8#36Artificial Analysis
Vals Indexnot in index57.0#17Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
Meta136 tok/s3.92 s$1.25$4.251.0M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.150 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0018$0.001425.1 s
Summarise a 30-page report12,000 / 600$0.018$0.007726.5 s
Code edit6,000 / 1,500$0.014$0.008930.6 s
Agentic coding session60,000 / 4,000$0.092$0.04342.0 s
Structured extraction2,000 / 200$0.0034$0.001724.6 s

See also

Data as of 9 Sept 2026. Compare these configurations.