BenchLeader
AmazonOpen weightsReasoning model

GLM-4.7-Flash

GLM-4.7-Flash is an Amazon open-weights reasoning model, released 19 Jan 2026. It is not yet ranked: no independent benchmark results so far, only pricing and metadata. At $0.153 per million tokens blended it is among the cheapest fifth of ranked models.

Blended price
$0.153/M
$0.070 in · $0.400 out
Output speed
First answer
Context
200k
Released
19 Jan 2026

No independent quality benchmark results yet — only pricing and metadata.

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0001
Summarise a 30-page report12,000 / 600$0.0011
Code edit6,000 / 1,500$0.0010
Agentic coding session60,000 / 4,000$0.0058
Structured extraction2,000 / 200$0.0002

See also

Data as of 13 Sept 2026. Compare with another model.

Cite as: BenchLeader, “GLM-4.7-Flash: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/zai-glm-4-7-flash, data as of 13 Sept 2026.