AA-Omniscience: non-hallucination
How often a model declines to answer rather than inventing one, on questions it does not know. The counterweight to the accuracy score.
As of 22 Sept 2026, MiniCPM5-1B leads AA-Omniscience: non-hallucination on BenchLeader with 99.1%, ahead of G9v3-3B at 88.3%, across 501 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 501
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1MiniCPM5-1B (none)99.1%
- 2G9v3-3B88.3%
- 3G9v3-39A5B87.0%
- 4Command A+85.8%
- 5LFM2.5-2.6B84.0%
- 6Grok 4.3 (medium)83.1%
- 7MiniCPM5-1B (thinking)83.0%
- 8Grok 4.3 (low)82.6%
- 9Grok 4.20 (thinking)82.6%
- 10Qwen3.8 27B (none)81.9%
- 11MiniMax M381.6%
- 12Quasar 438B78.6%
- 13MiniCPM5-2B78.1%
- 14Ling 3.0 Flash VL78.0%
- 15K Exaone 2.0 080377.4%
501 of 501
| # | |||
|---|---|---|---|
| 1 | 99.1% | 43.8 | |
| 2 | G9v3-3BAI9Starsopen | 88.3% | – |
| 3 | G9v3-39A5BAI9Starsopen | 87.0% | 52.8 |
| 4 | 85.8% | 51.1 | |
| 5 | 84.0% | 44.0 | |
| 6 | 83.1% | 59.5 | |
| 7 | 83.0% | 44.5 | |
| 8 | 82.6% | 57.6 | |
| 9 | 82.6% | 59.7 | |
| 10 | 81.9% | 50.6 | |
| 11 | 81.6% | 57.2 | |
| 12 | 78.6% | 56.8 | |
| 13 | 78.1% | 49.9 | |
| 14 | 78.0% | 56.4 | |
| 15 | 77.4% | 50.3 | |
| 16 | 76.0% | 65.5 | |
| 17 | 76.0% | 64.1 | |
| 18 | 75.6% | 55.2 | |
| 19 | 75.3% | 58.9 | |
| 20 | 74.6% | 54.5 | |
| 21 | 74.4% | 49.7 | |
| 22 | 74.4% | 60.8 | |
| 23 | 74.3% | 49.1 | |
| 24 | 74.2% | 57.9 | |
| 25 | 74.2% | 50.0 | |
| 26 | 74.2% | 57.2 | |
| 27 | 73.7% | 63.6 | |
| 28 | 73.7% | 45.5 | |
| 29 | 72.7% | 47.3 | |
| 30 | 72.4% | 63.6 | |
| 31 | 72.3% | 59.5 | |
| 32 | 71.6% | 58.7 | |
| 33 | 71.2% | 66.2 | |
| 34 | 70.9% | 53.8 | |
| 35 | New | 70.7% | 61.9 |
| 36 | 70.5% | 65.4 | |
| 37 | 70.3% | 57.3 | |
| 38 | 70.0% | 57.4 | |
| 39 | 70.0% | 59.6 | |
| 40 | 69.7% | 58.9 | |
| 41 | 69.6% | 39.3 | |
| 42 | 69.5% | 49.9 | |
| 43 | 69.5% | 54.6 | |
| 44 | 69.3% | 60.5 | |
| 45 | 69.1% | 45.0 | |
| 46 | 68.5% | 67.9 | |
| 47 | 68.1% | 55.5 | |
| 48 | 68.0% | 46.7 | |
| 49 | 67.6% | 42.1 | |
| 50 | New | 67.6% | – |
| 51 | 67.2% | 60.3 | |
| 52 | 67.1% | 69.2 | |
| 53 | 67.0% | 53.7 | |
| 54 | 66.7% | 63.7 | |
| 55 | 66.4% | 52.4 | |
| 56 | 66.3% | 49.3 | |
| 57 | 65.7% | 64.2 | |
| 58 | 65.6% | 55.2 | |
| 59 | 65.4% | 58.4 | |
| 60 | 64.7% | 59.4 |
Cite as: BenchLeader, “AA-Omniscience: non-hallucination leaderboard”, https://www.benchleader.com/benchmarks/aa_omniscience_non_hallucination, data as of 22 Sept 2026.
AA-Omniscience: non-hallucination: questions
- What does AA-Omniscience: non-hallucination measure?
- How often a model declines to answer rather than inventing one, on questions it does not know. The counterweight to the accuracy score. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads AA-Omniscience: non-hallucination?
- MiniCPM5-1B leads AA-Omniscience: non-hallucination with 99.1% as of 22 Sept 2026, ahead of G9v3-3B at 88.3%.
- How many models have AA-Omniscience: non-hallucination results?
- 501 model configurations have a AA-Omniscience: non-hallucination result on BenchLeader, all taken from Artificial Analysis.
- Who runs AA-Omniscience: non-hallucination and how often is it updated?
- AA-Omniscience: non-hallucination is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does AA-Omniscience: non-hallucination count toward the BenchLeader Index?
- No. AA-Omniscience: non-hallucination is shown for reference but left out of the composite index.