BenchLeader

Grok 4.7 vs Qwen3.8-Flash-Next

Verdict
  • Grok 4.7 (high) leads on quality: 65.8 vs 61.5.
  • Grok 4.7 (high) is stronger in composite, knowledge, reasoning.
  • Qwen3.8-Flash-Next is stronger in long context, agents & tools, coding, multimodal.
  • Qwen3.8-Flash-Next is 13× cheaper ($0.230 vs $3.00 per 1M blended).
MetricGrok 4.7 (high)Qwen3.8-Flash-Next
BenchLeader Index65.861.5
Composite score87.567.2
Knowledge score72.659.4
Long context score64.866.2
Reasoning score76.463.3
Agents & tools score65.8
Coding score71.2
Multimodal score64.3
Blended price $/M$3.00$0.230
Output speed55 tok/s61 tok/s
Time to first answer1.4 s35.5 s
Context window500k256k
LMArena WebDev1635
LMArena Agent-0.1
LiveBench76.2%
LiveBench Reasoning87.4%
LiveBench Coding72.5%
LiveBench Agentic Coding61.6%
LiveBench Mathematics85.8%
LiveBench Data Analysis74.2%
LiveBench Language74.6%
LiveBench Instruction Following77.1%
AA Intelligence Index46.339.8
AA-LCR77.0%79.7%
MMMU-Pro79.8%
AA-Omniscience30.9-9.7
GPQA Diamond (AA)92.3%
Humanity's Last Exam (AA)42.3%38.0%
SciCode (AA)57.8%50.6%
LegalBench85.1%
CritPt18.0%11.1%
GDPval (AA)59.7%55.6%
τ²-Bench Banking (AA)45.4%
Code Migration34.7%
Excel Modeling Benchmark54.9%
Harvey's Legal Agent Benchmark19.6%
Legal Research Bench39.9%
MedCode48.7%
MedScribe87.2%
ProgramBench0.0%
Public Benefits Bench68.5%
SAGE31.0%

Data as of 2026-09-21. Best configuration of each model; every score links to its source on the model pages.

Grok 4.7 vs Qwen3.8-Flash-Next: questions

Is Grok 4.7 better than Qwen3.8-Flash-Next?
Grok 4.7 (high) leads on quality: 65.8 vs 61.5. The BenchLeader Index combines every independent quality benchmark; Grok 4.7 (high) is ahead overall as of 2026-09-21, but check the category scores for your use.
Which is cheaper, Grok 4.7 or Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next is cheaper: $0.230 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.7 or Qwen3.8-Flash-Next?
Qwen3.8-Flash-Next streams faster: 61 against 55 output tokens per second.
Which has the larger context window?
Grok 4.7 accepts more context: 500k against 256k tokens.