BenchLeader

Gemini 4 Argon

Reasoning effort

Gemini 4 Argon is a Google proprietary reasoning model, released 30 Sept 2026. Its best configuration (high reasoning effort) ranks #5 of 760 on the BenchLeader Index at 70.0 ±3.9, in the top ten. The ± is the point: 42 other configurations score within that range, so they and this one cannot be told apart on quality alone — price and speed are what separate them. It scores highest in composite (89) and lowest in long context (65). At $4.00 per million tokens blended it is among the most expensive ranked models. It has been measured at 2 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 8 Oct 2026.

Blended price
$4.00/M
$2.00 in · $10.00 out
Output speed
–
First answer
–
Context
1M
Full answer
–
Cost per run
$1.99
one full Intelligence Index run
Released
30 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index70
  2. Reasoning82
  3. Coding72
  4. Agents & tools68
  5. Maths72
  6. Knowledge76
  7. Human preference73
  8. Long context65
  9. Composite89

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
highbest70.0#5––$0.0038Agents & tools 68 · Coding 72 · Composite 89 · Human preference 73 · Knowledge 76 · Long context 65 · Maths 72 · Reasoning 82
not stated–––––Agents & tools 83 · Maths 79

How its index has moved

2 Oct 2026 to 9 Oct 2026
667074

The index is recomputed from scratch every day, so a line moves when a new benchmark result lands, when a publisher revises a score, or when the models it is normalised against change. Early movement usually means the score is still settling.

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkhighnot statedSourceTrend
LMArena Hard Prompts1551#1–LMArena
Humanity's Last Exam (AA)not in index57.1%#5–Artificial Analysis
CritPt27.1%#31–Artificial Analysis
MysteryMechanismnot in index45.5%#6–Vals AI

Coding

Benchmarkhighnot statedSourceTrend
SciCode61.8%#4–SciCode
LMArena Coding1562#1–LMArena
LMArena WebDev1678#9–LMArena
SciCode (AA)not in index61.8%#4–Artificial Analysis
IOI100.0%#1–Vals AI
Code Migrationnot in index68.2%#2–Vals AI
ProgramBenchnot in index2.5%#5–Vals AI
Vibe Code Bench v1.1not in index91.9%#2–Vals AI
FrontierSWEnot in index55.0%#5–Epoch AI Benchmarking Hub
Terminal-Bench 4.0 (AA)not in index57.1%#6–Artificial Analysis

Agents & tools

Benchmarkhighnot statedSourceTrend
APEX-Agentseffort not stated–82.2%#1Mercor
LMArena Agent9.3#7–LMArena
GDPval-AA v2.1not in index56.3%#21–Artificial Analysis
CUA-benchnot in index4.8%#7–Vals AI
CyberBenchnot in index77.9%#2–Vals AI
Finance Agent v2not in index65.4%#1–Vals AI
Harvey's Legal Agent Benchmarknot in index19.6%#5–Vals AI
Legal Research Benchnot in index54.8%#4–Vals AI
SREBenchnot in index44.3%#3–Vals AI
Tax Agent Benchnot in index76.2%#2–Vals AI
Terminal-Bench 4.0 (Vals)not in index57.6%#5–Vals AI
Terminal-Bench Sciencenot in index44.3%#3–Vals AI
Vending-Bench 2not in indexeffort not stated–13718.2#3Andon Labs
AutomationBenchnot in index77.5%#1–Artificial Analysis
GDP.pdfnot in index21.8%#21–Artificial Analysis
AA-Briefcase v1.1not in index1488#27–Artificial Analysis

Maths

Benchmarkhighnot statedSourceTrend
ProofBencheffort not stated–99.0%#4Vals AI
LMArena Maths1528#1–LMArena

Knowledge

Benchmarkhighnot statedSourceTrend
AA-Omniscience42.4#9–Artificial Analysis
LegalBench88.3%#3–Vals AI
BioMysteryBenchnot in index76.3%#6–Vals AI
Excel Modeling Benchmarknot in index75.2%#4–Vals AI
MedCodenot in index58.8%#3–Vals AI
MedScribenot in index87.4%#15–Vals AI
Public Benefits Benchnot in index69.8%#5–Vals AI
AA-Omniscience: accuracynot in index49.9%#61–Artificial Analysis
AA-Omniscience: non-hallucinationnot in index84.9%#5–Artificial Analysis

Instruction following

Benchmarkhighnot statedSourceTrend
LMArena Instruction Followingnot in index1529#1–LMArena

Human preference

Benchmarkhighnot statedSourceTrend
LMArena Text1525#1–LMArena
LMArena Creative Writingnot in index1519#1–LMArena
LMArena Multi-turnnot in index1553#1–LMArena
LMArena Longer Queriesnot in index1545#1–LMArena

Multimodal

Benchmarkhighnot statedSourceTrend
SAGEnot in index53.6%#5–Vals AI
Blueprint-Bench 2not in indexeffort not stated–54.5%#1Epoch AI Benchmarking Hub

Long context

Benchmarkhighnot statedSourceTrend
AA-LCR79.7%#83–Artificial Analysis

Composite

Benchmarkhighnot statedSourceTrend
AA Intelligence Index v4.3.252.6#8–Artificial Analysis
Vals Indexnot in index68.9#1–Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0038––
Summarise a 30-page report12,000 / 600$0.030––
Code edit6,000 / 1,500$0.027––
Agentic coding session60,000 / 4,000$0.160––
Structured extraction2,000 / 200$0.0060––

See also

Data as of 11 Oct 2026. Compare these configurations.

Cite as: BenchLeader, “Gemini 4 Argon: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/gemini-4-argon, data as of 11 Oct 2026.