NVIDIA
The GPU maker also publishes open Nemotron models fine-tuned from Llama and its own architectures, mostly as showcases for its inference stack.
- Provider rank
- #11
- by best model
- Best index
- 58.3
- Nemotron 3 Ultra 550B A55B
- Ranked models
- 19
- 32 open weights
- Cheapest ranked
- $0.063/M
- Gemma 3 12B IT
Release timeline
Each model’s best configuration by release date. The top eight are labelled; hover any point for its name.
Next release
Over the last two years NVIDIA shipped 24 distinct models, one every 26 days on median. The latest arrived on 11 Aug 2026.
If the cadence holds, the next one lands between 26 Aug 2026 and 16 Sept 2026.
An estimate from past release dates only. It knows nothing about roadmaps or announcements.
Fastest ranked model: nemotron-3-nano-30b-a3b at 231 tok/s.
Head to head
Models
Every NVIDIA model with a result, one row per model; switch to all configurations to see each reasoning effort.
48 of 48
| # | ||||||||
|---|---|---|---|---|---|---|---|---|
| 1 | 58.3±7 | $1.00 | 156 | 17 s | 1M | 2026-06-04 | 2026-09-05 | |
| 2 | 55.4±6 | $0.205 | 35 | 51 s | 256k | 2026-04-02 | 2026-09-08 | |
| 3 | 53.6±5 | $0.350 | 96 | 25 s | 262k | 2026-03-11 | 2026-09-02 | |
| 4 | 51.0±12 | – | – | – | 1M | – | – | |
| 5 | 48.6±10 | $0.088 | 231 | 9.82 s | 1M | 2025-12-15 | – | |
| 6 | 48.2±9 | – | – | – | – | 2026-06-04 | 2026-06-04 | |
| 7 | 48.0±5 | $0.300 | 26 | 80 s | 262k | 2025-08-20 | 2025-08-20 · stale | |
| 8 | 46.0±7 | $0.158 | – | – | 256k | 2026-04-28 | – | |
| 9 | 44.9±2 | – | – | – | – | 2026-03-11 | 2026-03-11 · stale | |
| 10 | 44.7±6 | $0.400 | – | – | 131k | 2025-07-25 | 2026-09-02 | |
| 11 | 44.1±8 | – | – | – | 262k | – | – | |
| 12 | 42.8±4 | $1.20 | 43 | 8.27 s | 128k | 2025-04-15 | 2026-09-02 | |
| 13 | 42.6±2 | $0.400 | 35 | 12 s | 128k | 2025-07-25 | – | |
| 14 | 42.0±3 | $0.300 | 42 | 58 s | 128k | 2025-10-28 | – | |
| 15 | 42.0±4 | $0.070 | 102 | 27 s | 131k | 2024-12-01 | – | |
| 16 | 37.2±5 | $0.063 | 5 | 1.17 s | 131k | 2025-03-12 | 2026-09-02 | |
| 17 | 37.1±4 | – | – | – | 128k | 2025-06-12 | – | |
| 18 | 35.5±9 | $0.131 | 45 | 0.84 s | 131k | 2025-02-27 | 2025-02-27 · stale | |
| 19 | 34.5±5 | $0.063 | 27 | 0.62 s | 131k | 2025-03-12 | 2026-09-02 | |
| 20 | – | $0.400 | 58 | 44 s | 128k | 2025-07-25 | – | |
| 21 | – | – | – | – | 128k | 2025-04-07 | – | |
| 22 | – | – | – | – | 128k | 2024-07-16 | 2026-09-02 | |
| 23 | – | – | – | – | – | – | 2026-09-02 | |
| 24 | – | – | – | – | – | – | 2026-09-02 | |
| 25 | – | – | – | – | – | – | 2026-09-02 | |
| 26 | – | – | – | – | – | – | 2026-09-02 | |
| 27 | – | $0.095 | 277 | 7.89 s | 1M | 2026-08-11 | – | |
| 28 | – | – | – | – | – | – | 2026-09-02 | |
| 29 | – | – | – | – | – | – | 2026-09-02 | |
| 30 | – | $0.140 | 17 | 0.84 s | 128k | 2024-12-11 | – | |
| 31 | – | – | – | – | 128k | 2024-06-05 | 2026-09-02 | |
| 32 | – | – | – | – | – | – | 2026-09-05 | |
| 33 | – | – | – | – | – | – | 2026-08-27 | |
| 34 | – | – | – | – | 8k | 2024-01-30 | – | |
| 35 | – | – | – | – | 131k | 2025-12-01 | – | |
| 36 | – | – | – | – | 128k | 2024-09-11 | – | |
| 37 | – | – | – | – | 128k | 2026-03-03 | – | |
| 38 | – | $0.075 | 118 | 0.18 s | 262k | 2026-07-02 | – | |
| 39 | – | – | – | – | 131k | 2025-03-18 | – | |
| 40 | – | – | – | – | 33k | 2025-04-10 | – | |
| 41 | – | – | – | – | 262k | 2026-08-11 | – | |
| 42 | – | – | – | – | – | 2024-02-26 | – | |
| 43 | – | – | – | – | 128k | 2024-08-21 | – | |
| 44 | – | – | – | – | 128k | 2026-03-16 | – | |
| 45 | – | – | – | – | 128k | 2025-12-12 | – | |
| 46 | – | – | – | – | – | – | – | |
| 47 | – | – | – | – | – | – | – | |
| 48 | – | – | – | – | – | – | – |
Latest news
- Nvidia bets $13 billion on open AI models with Hugging Face dealReuters · 4 Sept 2026
- NVIDIA and Local AI Community Fuel Open Source Models and Intelligent AgentsNVIDIA Blog · 11 Aug 2026
- Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference | NVIDIA Technical BlogNVIDIA Developer · 31 Jul 2026
- Nvidia and CrowdStrike Develop New Cybersecurity AI Models | The Morning Download for Sept. 2WSJ · 9 Sept 2026
- Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference | NVIDIA Technical BlogNVIDIA Developer · 2 Sept 2026
- Versos AI Launches Agents Built with NVIDIA NeMo to Speed Curation of Video Training Data for AI Modelsaithority.com · 9 Sept 2026
- Nvidia’s Hugging Face Deal Is a Bid for A.I.’s Next Control Pointobserver.com · 9 Sept 2026
- NVIDIA Bets Big on Hugging Face: ETFs to WinTradingView · 9 Sept 2026