BenchLeader

AA-Omniscience: accuracy

The share of AA-Omniscience questions answered correctly, before any penalty for confident wrong answers.

As of 22 Sept 2026, Claude Fable 5.1 leads AA-Omniscience: accuracy on BenchLeader with 67.2%, ahead of Claude Fable 5.1 at 66.2%, across 501 model configurations with a published result.

Published by
Artificial Analysis
Category
Knowledge
Index weight
Reference only
Models
501
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

501 of 501
#
1Claude Fable 5.1thinkingAnthropic67.2%71.0
2Claude Fable 5.1xhighAnthropic66.2%71.1
3NewClaude Opus 5.5thinkingAnthropic66.2%70.7
4NewClaude Opus 5.5xhighAnthropic65.4%70.3
5Claude Fable 5thinkingAnthropic65.3%70.3
6Claude Fable 5.1highAnthropic64.9%71.0
7NewClaude Opus 5.5highAnthropic64.6%69.9
8NewClaude Opus 5.5mediumAnthropic64.5%69.3
9NewClaude Opus 5.5lowAnthropic63.5%65.2
10Claude Fable 5.1mediumAnthropic63.1%69.1
11GPT-6 AstramaxOpenAI62.6%71.7
12GPT-6 AstraxhighOpenAI61.9%70.5
13GPT-6 AstrahighOpenAI61.1%71.3
14Claude Opus 5maxAnthropic60.9%69.5
15GPT-6 AstramediumOpenAI60.6%69.2
16Claude Fable 5.1lowAnthropic60.2%66.7
17GPT-6 AstralowOpenAI59.5%67.6
18Claude Opus 5xhighAnthropic59.5%69.3
19GPT-5.6 SolmaxOpenAI59.4%68.5
20Claude Opus 5highAnthropic58.9%69.8
21GPT-5.6 SolxhighOpenAI58.8%67.9
22GPT-5.6 SolhighOpenAI58.4%67.4
23GPT-5.5xhighOpenAI58.0%67.2
24GPT-5.6 SolmediumOpenAI57.8%65.5
25GPT-5.6 SollowOpenAI57.2%63.2
26Claude Opus 5mediumAnthropic57.1%66.4
27GPT-5.5highOpenAI57.0%66.6
28GPT-5.5mediumOpenAI56.8%64.0
29Claude Opus 5lowAnthropic56.0%62.8
30Gemini 3 ProhighGoogle55.8%60.7
31Gemini 3.7 FlashhighGoogle55.3%64.3
32Gemini 3.1 ProGoogle54.9%63.2
33GPT-5.5lowOpenAI54.9%60.8
34Gemini 3.8 FlashhighGoogle54.6%64.3
35Gemini 3.7 FlashmediumGoogle54.0%64.4
36Gemini 3.7 FlashlowGoogle53.6%61.6
37Gemini 3 FlashthinkingGoogle53.4%60.0
38Gemini 3.8 FlashmediumGoogle53.0%64.6
39GPT-5.3-CodexxhighOpenAI52.9%64.9
40Gemini 3.8 FlashlowGoogle52.2%60.2
41Muse Spark 1.1xhighMeta52.0%61.4
42Grok 4.5highSpaceXAI51.5%60.6
43Grok Build 0.1SpaceXAI51.5%54.5
44Gemini 3.5 FlashhighGoogle51.4%63.4
45Gemini 3.5 FlashmediumGoogle51.0%63.9
46GPT-5.4xhighOpenAI50.9%65.1
47Gemini 3.6 FlashhighGoogle50.0%61.5
48Muse SparkMeta49.6%65.4
49DeepSeek V4 PromaxDeepSeekopen ↗49.1%63.7
50Claude Opus 4.7maxAnthropic48.9%63.7
51Claude Opus 4.8maxAnthropic48.8%63.5
52GPT-5.6 Solno reasoningOpenAI48.7%54.8
53Gemini 3 ProlowGoogle48.3%54.4
54Grok 4.6highSpaceXAI48.2%64.2
55GPT-5.4lowOpenAI47.9%60.5
56NewGrok 4.7highSpaceXAI47.8%
57Kimi K3maxMoonshot AIopen ↗47.6%66.7
58NewGrok 4.7xhighSpaceXAI47.5%61.9
59Claude Opus 4.6maxAnthropic47.0%58.5
60GPT-5.6 TerramaxOpenAI46.8%64.8

Cite as: BenchLeader, “AA-Omniscience: accuracy leaderboard”, https://www.benchleader.com/benchmarks/aa_omniscience_accuracy, data as of 22 Sept 2026.

AA-Omniscience: accuracy: questions

What does AA-Omniscience: accuracy measure?
The share of AA-Omniscience questions answered correctly, before any penalty for confident wrong answers. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads AA-Omniscience: accuracy?
Claude Fable 5.1 leads AA-Omniscience: accuracy with 67.2% as of 22 Sept 2026, ahead of Claude Fable 5.1 at 66.2%.
How many models have AA-Omniscience: accuracy results?
501 model configurations have a AA-Omniscience: accuracy result on BenchLeader, all taken from Artificial Analysis.
Who runs AA-Omniscience: accuracy and how often is it updated?
AA-Omniscience: accuracy is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does AA-Omniscience: accuracy count toward the BenchLeader Index?
No. AA-Omniscience: accuracy is shown for reference but left out of the composite index.