Llama 3.3 70B
Best configuration ranks #500 of 610 on the BenchLeader Index at 40.7 ±4.3. Last measured 2 Sept 2026. Released 6 Dec 2024.
- Blended price
- $0.620/M
- $0.590 in · $0.710 out
- Output speed
- 86 tok/s
- First answer
- 1.64 s
- first token 0.75 s
- Context
- 131k
- Overall index41
- Reasoning35
- Coding33
- Agents & tools42
- Maths34
- Knowledge38
- Instruction following49
- Human preference48
- Long context33
- Composite39
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Reasoning
| Benchmark | default | Source |
|---|---|---|
| GPQA Diamond | 47.4%#210 | Epoch AI Benchmarking Hub |
| SimpleBench | 19.9%#79 | SimpleBench |
| LMArena Hard Prompts | 1320#221 | LMArena |
| GPQA Diamond (AA)not in index | 49.8%#408 | Artificial Analysis |
| Humanity's Last Exam (AA)not in index | 3.6%#502 | Artificial Analysis |
Coding
Agents & tools
| Benchmark | default | Source |
|---|---|---|
| Terminal-Bench Hard | 3.0%#298 | Artificial Analysis |
| τ²-Bench Telecom (AA)not in index | 26.6%#271 | Artificial Analysis |
| BFCL Overall | 31.9%#46 | Berkeley Function Calling Leaderboard |
Maths
| Benchmark | default | Source |
|---|---|---|
| OTIS Mock AIME | 5.1%#225 | Epoch AI Benchmarking Hub |
| MATH Level 5 | 41.6%#58 | Epoch AI Benchmarking Hub |
Knowledge
| Benchmark | default | Source |
|---|---|---|
| AA-Omniscience | -54.2#352 | Artificial Analysis |
Instruction following
| Benchmark | default | Source |
|---|---|---|
| IFBench | 47.1%#188 | Artificial Analysis |
Human preference
| Benchmark | default | Source |
|---|---|---|
| LMArena Text | 1317#216 | LMArena |
Long context
| Benchmark | default | Source |
|---|---|---|
| AA-LCR | 15.7%#379 | Artificial Analysis |
Composite
| Benchmark | default | Source |
|---|---|---|
| Epoch Capabilities Indexnot in index | 127.3#142 | Epoch AI Benchmarking Hub |
| AA Intelligence Index | 7.7#385 | Artificial Analysis |
Where to run it
Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.
| Provider | Speed | First token | Input $/M | Output $/M | Context | Quantisation |
|---|---|---|---|---|---|---|
| Groq | 120 tok/s | 0.25 s | $0.590 | $0.790 | 131k | – |
| CoreWeave | 79 tok/s | 0.80 s | $0.710 | $0.710 | 128k | fp16 |
| SambaNova Turbo | 55 tok/s | 0.86 s | $0.450 | $0.900 | 131k | – |
| Crusoe | 52 tok/s | 0.75 s | $0.250 | $0.750 | 131k | bf16 |
| Parasail | 42 tok/s | 0.61 s | $0.220 | $0.500 | 131k | fp8 |
| Google Vertex (US) | 42 tok/s | 0.32 s | $0.720 | $0.720 | 128k | – |
| Google Vertex | 34 tok/s | 0.33 s | $0.720 | $0.720 | 128k | – |
| Cloudflare | 31 tok/s | 0.51 s | $0.293 | $2.25 | 24k | fp8 |
| NovitaAI | 30 tok/s | 1.13 s | $0.135 | $0.400 | 12k | bf16 |
| AkashML | 28 tok/s | 1.36 s | $0.200 | $0.520 | 131k | fp8 |
| Together | 26 tok/s | 0.80 s | $1.04 | $1.04 | 131k | – |
| DeepInfra (Turbo) | 13 tok/s | 0.47 s | $0.100 | $0.320 | 131k | fp8 |
| Nebius Token Factory | 4 tok/s | 5.25 s | $0.130 | $0.400 | 131k | fp8 |
Price history
Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.
What a task costs
Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.
| Workload | Tokens in / out | Cost | With caching | Time |
|---|---|---|---|---|
| Chat reply | 400 / 300 | $0.0004 | – | 5.1 s |
| Summarise a 30-page report | 12,000 / 600 | $0.0075 | – | 8.6 s |
| Code edit | 6,000 / 1,500 | $0.0046 | – | 19.1 s |
| Agentic coding session | 60,000 / 4,000 | $0.038 | – | 48.1 s |
| Structured extraction | 2,000 / 200 | $0.0013 | – | 4.0 s |
See also
Data as of 9 Sept 2026. Compare with another model.