BenchLeader

Open weights vs closed models in 2026: how far behind, and where it doesn't matter

The best open-weights LLMs ranked against the closed frontier, with the gap in points, price and speed, from independent benchmarks.

You can download the weights of a model that would have topped every leaderboard eighteen months ago. Whether that is good enough depends on the job, so this page puts a number on the gap rather than an opinion. Data as of 9 Sept 2026.

The best open-weights models today

#ModelIndexIndexPrice $/M
1 Kimi K3 Moonshot AI66.566.5$6.00
2 GLM-5.3-Flash Zhipu AI61.061.0$0.119
3 Kimi K2.6 Moonshot AI60.660.6$1.71
4 GLM-5.1 Zhipu AI60.160.1$2.15
5 DeepSeek V4 Pro high DeepSeek60.060.0$0.548

As of 9 Sept 2026. Full list.

The overall leader is GPT-6 Astra (max); the best open-weights configuration, Kimi K3, is the first row above. The difference between those two numbers is the gap, and it has been closing at a steady few points a year.

Where the gap is small

Knowledge, maths competitions and instruction following: open models trained on the same public data and distilled from the same techniques land close to the frontier. On LiveBench and MMLU-Pro the top open configuration is usually within a handful of points of the best closed one.

Where it is large

Long-horizon agentic work: Terminal-Bench, SWE-Bench Pro, τ²-bench. These reward reliability over dozens of steps, and the frontier labs' post-training still shows. Humanity's Last Exam and FrontierMath Tier 4 also favour the closed frontier, where the last few points of reasoning are hardest to buy.

Price is the argument

Open weights are served by many hosts, and hosts compete. The "where to run it" table on each model page shows the same model at prices that differ by 3× between providers, with measured speed for each. Against that, the closed frontier's list price is the list price. On the value board, open models occupy most of the top.

Speed is a wash, or better

Hosts specialising in fast inference serve open models at hundreds of tokens per second. Several of the fast and good entries are open weights.

What "open" means here

We mark a model open when its weights are publicly downloadable, and we link the official repository from the model page wherever a source names it. Licences vary from permissive to restricted commercial use; we do not summarise them, so read the licence on the repository before you build on it.

How to decide

If your work looks like a chat assistant, a summariser or a structured-extraction pipeline, the gap will not show up in your product and the price difference will. If it looks like an autonomous agent running for an hour, run both on your own tasks before choosing; the benchmarks above say the frontier still earns its premium there, but the premium is shrinking every quarter.