The Lab Notes

Z.ai

Z.ai's faster GLM-5.3-Flash tier costs 2.5 times as much

Z.ai turned on GLM-5.3-FlashX, a quicker-serving version of GLM-5.3-Flash priced well above the standard tier and left off the flat-rate Coding Plan.

A speedometer needle pushed into the red beside a stack of coins.

Z.ai has turned on a faster way to serve its GLM-5.3-Flash model, and developers who want the speed will pay roughly 2.5 times the standard rate for it rather than get a new model.

The tier, called GLM-5.3-FlashX, is live and delivers inference at about 200 tokens per second, according to Z.ai’s developer documentation, confirmed live Sept. 21, 2026, which states plainly that “GLM-5.3-FlashX is now live, delivering inference speeds of 200 tokens/s.” Vercel’s AI Gateway changelog, dated Sept. 18, 2026, recorded the same speed when it added FlashX to the platform.

FlashX is not a new model. It runs on the identical underlying weights as standard GLM-5.3-Flash, Z.ai’s documentation said: 320 billion total parameters with 18 billion active, a hybrid sparse-and-linear attention architecture, and a 1 million-token context window. FlashX is a faster serving configuration of those same weights, not a separate release.

That speed carries a markup. Z.ai prices GLM-5.3-FlashX at $0.37 per million input tokens and $1.25 per million output tokens, with cached input at $0.075 per million tokens, the documentation showed. Standard GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens, with cached input at $0.03 per million tokens, roughly 2.5 times less than the faster tier.

The two tiers also sit on different access paths. Z.ai’s documentation said GLM-5.3-FlashX “is not yet available on” the GLM Coding Plan, the company’s flat-rate subscription, while standard GLM-5.3-Flash “is now fully available on the GLM Coding Plan.” For now, reaching FlashX means paying per token through the API or a third-party gateway such as Vercel’s, not the flat monthly plan.

Z.ai has not published benchmark scores specific to FlashX.

Analysis

Charging more to serve the same weights faster is an ordinary trade. Leaving that faster tier off the flat-rate Coding Plan is the odder call: a subscriber paying Z.ai’s monthly fee for coding help gets the slower Flash by default, and the tier built for latency-sensitive work is reachable only by switching to metered billing.