← Founder Notes
Archive

The model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on…

Yethikrishna ROriginal on Threads

the model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on hybrid sparse-plus-linear attention, hits 200 tokens a second, and cuts kv cache 4.4x.

inference speed is the release axis now, not leaderboard points.

Context

Z.ai's post of 26 August 2026 describes GLM-5.3-Flash as 320B total and 18B active parameters on hybrid sparse and linear attention, with 3.0x less attention compute and a 4.4x smaller KV cache than GLM-5.3, a KV cache it says is still slightly larger than Kimi-K3 and DeepSeek-V4-Flash, with weights on Hugging Face. Z.ai claims it outperforms GLM-5.2 at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks, which are vendor claims.

How it compares

The architecture facts are first-party for GLM-5.3-Flash, not for glm-5.3-flashx. The 4.4x is against the larger GLM-5.3, not against Flash or FlashX. The 200 tokens a second and the FlashX designation were not on the Z.ai pages read, and the pricing page copy has no GLM-5.3 FlashX row. The earlier FlashX note drew on secondary reports of a launch on 18 September 2026 at a vendor-peak 200 tokens a second, which stay as positive secondary evidence and are not reconciled with a first-party FlashX page. The Flash post leads with benchmarks and price as well as efficiency. That inference speed is the release axis is the author's thesis.

Related work

Watch next

  • A first-party FlashX page with its spec and tokens-per-second configuration, and independent throughput tests.

Sources

  1. GLM-5.3-Flash (Z.ai, 26 Aug 2026)z.ai
  2. GLM-5.3-Flash guide (Z.ai docs)docs.z.ai
  3. Pricing (Z.ai docs)docs.z.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 04:04 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-model-race-now-has-a-speed-tier-DdfGGF5kRgb" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on…"></iframe>

More notes