BenchLeader

Ling 3.0 Flash

Best configuration ranks #210 of 610 on the BenchLeader Index at 54.0 ±4.3. Last measured 5 Sept 2026. Released 4 Aug 2026.

Blended price
$0.115/M
$0.080 in · $0.220 out
Output speed
290 tok/s
First answer
9.45 s
first token 0.65 s
Context
262k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index54
  2. Coding61
  3. Agents & tools39
  4. Knowledge53
  5. Long context63
  6. Composite60

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index85.5%#124Artificial Analysis
Humanity's Last Exam (AA)not in index23.7%#141Artificial Analysis
GPQA Diamond (Vals)not in index84.8%#48Vals AI

Coding

BenchmarkdefaultSource
SciCode (AA)not in index42.0%#107Artificial Analysis
LiveCodeBench84.0%#42Vals AI
SWE-bench (Vals)not in index65.2%#70Vals AI

Agents & tools

BenchmarkdefaultSource
Terminal-Bench 2.1 (Vals)50.2%#48Vals AI

Knowledge

BenchmarkdefaultSource
AA-Omniscience-17.9#176Artificial Analysis
MMLU-Pro82.0%#77Vals AI
LegalBench79.7%#86Vals AI
CorpFin61.9%#50Vals AI
TaxEval70.7%#86Vals AI

Long context

BenchmarkdefaultSource
AA-LCR73.0%#132Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index24.9#116Artificial Analysis
Vals Indexnot in index21.7#48Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
DeepInfra115 tok/s0.55 s$0.060$0.180131kbf16
NovitaAI15 tok/s0.76 s$0.021$0.063262k

Price history

Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.

$0.00$0.061$0.121$0.182$0.242Aug 26Sept 26
input outputnow $0.060 in · $0.180 out

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.000110.5 s
Summarise a 30-page report12,000 / 600$0.001111.5 s
Code edit6,000 / 1,500$0.000814.6 s
Agentic coding session60,000 / 4,000$0.005723.2 s
Structured extraction2,000 / 200$0.000210.1 s

See also

Data as of 9 Sept 2026. Compare with another model.