BenchLeader

CommonsenseQA 2

Commonsense yes/no questions. A 2022 benchmark, saturated for current models; kept for history.

As of 19 Sept 2026, T5-11B leads CommonsenseQA 2 on BenchLeader with 67.8%, ahead of T5-3B at 60.2%, across 6 model configurations with a published result.

Published by
Epoch AI Benchmarking Hub
Category
Reasoning
Index weight
Reference only
Models
6
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Commonsense yes/no questions.

How it is scored

Accuracy, as compiled by Epoch AI from published results.

What to keep in mind

Saturated: frontier models score near the ceiling, so it separates only older or smaller models, and results are largely self-reported by developers.

6 of 6
#
1T5-11BGoogle67.8%
2T5-3BGoogle60.2%
3GPT-3.5 Turbo (older v0613)OpenAI57.0%
4T5-LargeUnknown54.6%
5InstructGPT 175BOpenAI52.9%
6Llama 2-70BMetaopen ↗50.0%34.9

Cite as: BenchLeader, “CommonsenseQA 2 leaderboard”, https://www.benchleader.com/benchmarks/csqa2, data as of 19 Sept 2026.

CommonsenseQA 2: questions

What does CommonsenseQA 2 measure?
Commonsense yes/no questions. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads CommonsenseQA 2?
T5-11B leads CommonsenseQA 2 with 67.8% as of 19 Sept 2026, ahead of T5-3B at 60.2%.
How many models have CommonsenseQA 2 results?
6 model configurations have a CommonsenseQA 2 result on BenchLeader, all taken from Epoch AI Benchmarking Hub.
Who runs CommonsenseQA 2 and how often is it updated?
CommonsenseQA 2 is published by Epoch AI Benchmarking Hub. BenchLeader re-reads the published results every morning and records the date each result was published.
Does CommonsenseQA 2 count toward the BenchLeader Index?
No. CommonsenseQA 2 is shown for reference but left out of the composite index.