BenchLeader

Ling 3.1 Flash

Ling 3.1 Flash is an Ant Group proprietary reasoning model, released 29 Sept 2026. It is not yet ranked: 4 independent results so far, and the index needs at least five across two categories. At $0.450 per million tokens blended it is cheaper than most ranked models. Output speed of 214 tokens per second puts it in the fastest quarter, with a first answer in 11.1 s.

Blended price
$0.450/M
$0.300 in · $0.900 out
Output speed
214 tok/s
measured by Artificial Analysis
First answer
11 s
first token 2.51 s
Context
262k
Full answer
13 s
median, reasoning included
Cost per run
$0.986
one full Intelligence Index run
Released
29 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Reasoning71
  2. Knowledge63
  3. Long context67
  4. Composite76

Versions

Ant Group has shipped 3 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
Ling 3.1 Flashthis page29 Sept 2026––
Ling 3.0 Flash23 Jul 202652.9#284
Ling Flash 2.018 Sept 202544.6#510

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarknot statedSource
Humanity's Last Exam (AA)not in index39.4%#85Artificial Analysis
CritPt18.0%#66Artificial Analysis

Coding

Benchmarknot statedSource
SciCode (AA)not in index54.0%#64Artificial Analysis
Terminal-Bench 4.0 (AA)not in index33.3%#35Artificial Analysis

Agents & tools

Benchmarknot statedSource
GDPval-AA v2.1not in index56.1%#22Artificial Analysis

Long context

Benchmarknot statedSource
AA-LCR83.0%#24Artificial Analysis

Composite

Benchmarknot statedSource
AA Intelligence Index v4.3.241.1#54Artificial Analysis

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
NovitaAI69 tok/s2.71 s$0.000$0.000262k–

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0004–12.5 s
Summarise a 30-page report12,000 / 600$0.0041–13.9 s
Code edit6,000 / 1,500$0.0032–18.2 s
Agentic coding session60,000 / 4,000$0.022–29.9 s
Structured extraction2,000 / 200$0.0008–12.1 s

See also

Data as of 11 Oct 2026. Compare with another model.

Cite as: BenchLeader, “Ling 3.1 Flash: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/ling-3-1-flash, data as of 11 Oct 2026.