GLM-5.3-Flash API pricing: 30 providers compared

30 providers serve GLM-5.3-Flash. The cheapest charges $0.07 per million input tokens, the most expensive $0.45. Put in your monthly usage to see what that gap costs you.

Provider$/M in$/M out$/M cached inYour monthPrecisionContextUptime 1dTok/s

Shipping with AI? New models, SDK releases and price changes that matter, checked against the sources. One email a day.

Precision is what each provider reports. Some models ship in fp8 or fp4, so a low number alone does not mean a cut-down copy. If a provider runs lower precision than the release, its answers can differ from the version you tested. "unknown" means the provider does not say.

About GLM-5.3-Flash

Sam Witteveen puts Z.ai's GLM-5.3-Flash next to the full GLM-5.3 and asks when the cheaper model is the better pick. Flash is 320B parameters with 18B active; GLM-5.3 is 753B.

What changed in this release · Model card on Hugging Face

Where these prices come from

Prices, precision, context, uptime and throughput are read live from OpenRouter's public endpoint list for z-ai/glm-5.3-flash each time this page loads, so they match what the providers charge today. If that request fails, the page shows the snapshot taken when it was built, and says so. Prices are per million tokens in US dollars, before any provider discounts or free tiers.