BenchLeader

AA-Omniscience: non-hallucination

How often a model declines to answer rather than inventing one, on questions it does not know. The counterweight to the accuracy score.

As of 22 Sept 2026, MiniCPM5-1B leads AA-Omniscience: non-hallucination on BenchLeader with 99.1%, ahead of G9v3-3B at 88.3%, across 501 model configurations with a published result.

Published by
Artificial Analysis
Category
Knowledge
Index weight
Reference only
Models
501
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

501 of 501
#
1MiniCPM5-1Bno reasoningOpenBMBopen99.1%43.8
2G9v3-3BAI9Starsopen88.3%
3G9v3-39A5BAI9Starsopen87.0%52.8
4Command A+Cohereopen85.8%51.1
5LFM2.5-2.6BLiquid AIopen ↗84.0%44.0
6Grok 4.3mediumSpaceXAI83.1%59.5
7MiniCPM5-1BthinkingOpenBMBopen83.0%44.5
8Grok 4.3lowSpaceXAI82.6%57.6
9Grok 4.20thinkingSpaceXAI82.6%59.7
10Qwen3.8 27Bno reasoningAlibabaopen ↗81.9%50.6
11MiniMax M3MiniMaxopen ↗81.6%57.2
12Quasar 438BMultiverse Computing78.6%56.8
13MiniCPM5-2BOpenBMBopen78.1%49.9
14Ling 3.0 Flash VLAnt Groupopen ↗78.0%56.4
15K Exaone 2.0 0803LG AI Researchopen77.4%50.3
16Grok 4.6mediumSpaceXAI76.0%65.5
17Grok 4.6xhighSpaceXAI76.0%64.1
18Solar Pro 4Upstage75.6%55.2
19MiMo-V2.5-ProXiaomiopen ↗75.3%58.9
20Solar Open2 250BUpstageopen ↗74.6%54.5
21Granite 4.2 30BIBMopen ↗74.4%49.7
22Qwen3 7maxAlibaba74.4%60.8
23Claude Haiku 4.5no reasoningAnthropic74.3%49.1
24Grok 4.3highSpaceXAI74.2%57.9
25Grok 3 minithinkingSpaceXAI74.2%50.0
26K2 Horizon 375B A23BMBZUAIopen74.2%57.2
27GLM 5.2maxZhipu AIopen ↗73.7%63.6
28Granite 4.2 3BIBMopen ↗73.7%45.5
29Claude Haiku 4.5thinkingAnthropic72.7%47.3
30GLM 5.3 FlashZhipu AIopen ↗72.4%63.6
31Qwen3.7 PlusAlibaba72.3%59.5
32Motif 3Motif Technologiesopen71.6%58.7
33Qwen3.8 Max (0902)maxAlibaba71.2%66.2
34Claude Sonnet 4thinkingAnthropic70.9%53.8
35NewGrok 4.7xhighSpaceXAI70.7%61.9
36GLM 5.3maxZhipu AIopen ↗70.5%65.4
37Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗70.3%57.3
38GLM 5.1thinkingZhipu AIopen ↗70.0%57.4
39Mimo v2 ProXiaomi70.0%59.6
40Qwen3.8 27BxhighAlibabaopen ↗69.7%58.9
41Gemma 3 270MGoogleopen ↗69.6%39.3
42Ling 3.0 TinyAnt Groupopen69.5%49.9
43K2 Horizon MoVA 36B A4BMBZUAIopen69.5%54.6
44Grok 4.6lowSpaceXAI69.3%60.5
45Gemma 4 E4BthinkingGoogleopen ↗69.1%45.0
46Muse Spark 1.3xhighMeta68.5%67.9
47MiMo-V2.5Xiaomiopen ↗68.1%55.5
48Granite 4.2 8BIBMopen ↗68.0%46.7
49Gemma 4 E2BthinkingGoogleopen ↗67.6%42.1
50NewGrok 4.7highSpaceXAI67.6%
51Agnes 2.5 Pro BetaSapiens AI67.2%60.3
52Muse Spark 1.3maxMeta67.1%69.2
53A.X-K2SK Telecomopen67.0%53.7
54Muse Spark 1.2xhighMeta66.7%63.7
55K2 Horizon 7BMBZUAIopen66.4%52.4
56GLM 5.2no reasoningZhipu AIopen ↗66.3%49.3
57Grok 4.6highSpaceXAI65.7%64.2
58Gemini 3.5 Flash LiteGoogle65.6%55.2
59Qwen3.6 PlusAlibaba65.4%58.4
60GLM-5thinkingZhipu AIopen ↗64.7%59.4

Cite as: BenchLeader, “AA-Omniscience: non-hallucination leaderboard”, https://www.benchleader.com/benchmarks/aa_omniscience_non_hallucination, data as of 22 Sept 2026.

AA-Omniscience: non-hallucination: questions

What does AA-Omniscience: non-hallucination measure?
How often a model declines to answer rather than inventing one, on questions it does not know. The counterweight to the accuracy score. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads AA-Omniscience: non-hallucination?
MiniCPM5-1B leads AA-Omniscience: non-hallucination with 99.1% as of 22 Sept 2026, ahead of G9v3-3B at 88.3%.
How many models have AA-Omniscience: non-hallucination results?
501 model configurations have a AA-Omniscience: non-hallucination result on BenchLeader, all taken from Artificial Analysis.
Who runs AA-Omniscience: non-hallucination and how often is it updated?
AA-Omniscience: non-hallucination is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does AA-Omniscience: non-hallucination count toward the BenchLeader Index?
No. AA-Omniscience: non-hallucination is shown for reference but left out of the composite index.