BenchLeader
ApodexReasoning modelAuto-detected

Apodex 1.1

Not yet ranked: too few independent results so far.

Blended price
$0.975/M
$0.300 in · $3.00 out
Output speed
First answer
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Knowledge54
  2. Multimodal64
  3. Long context66
  4. Composite67

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkdefaultSource
GPQA Diamond (AA)not in index86.4%#111Artificial Analysis
Humanity's Last Exam (AA)not in index34.1%#84Artificial Analysis

Coding

BenchmarkdefaultSource
SciCode (AA)not in index45.5%#90Artificial Analysis

Knowledge

BenchmarkdefaultSource
AA-Omniscience-21.9#186Artificial Analysis

Multimodal

BenchmarkdefaultSource
MMMU-Pro79.2%#43Artificial Analysis

Long context

BenchmarkdefaultSource
AA-LCR79.3%#65Artificial Analysis

Composite

BenchmarkdefaultSource
AA Intelligence Index30.4#79Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0010
Summarise a 30-page report12,000 / 600$0.0054
Code edit6,000 / 1,500$0.0063
Agentic coding session60,000 / 4,000$0.030
Structured extraction2,000 / 200$0.0012

See also

Data as of 9 Sept 2026. Compare with another model.