BenchLeader

LFM2.5-2.6B

Not yet ranked: too few independent results so far. Released 4 Aug 2026.

Blended price
Output speed
199 tok/s
First answer
14 s
Context
128k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Knowledge59
  2. Long context28
  3. Composite40

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index55.8%#381Artificial Analysis
Humanity's Last Exam (AA)not in index6.2%#327Artificial Analysis

Coding

BenchmarkdefaultSource
SciCode (AA)not in index14.3%#162Artificial Analysis

Knowledge

BenchmarkdefaultSource
AA-Omniscience-10.9#148Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR5.7%#413Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index8.4#361Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 30015.1 s
Summarise a 30-page report12,000 / 60016.6 s
Code edit6,000 / 1,50021.1 s
Agentic coding session60,000 / 4,00033.6 s
Structured extraction2,000 / 20014.6 s

See also

Data as of 9 Sept 2026. Compare with another model.