Z.AI: GLM 5.3 Flash

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Model specifications

Input
text, image, video
Output
text
Context
1,000,000 tokens
Max output
131,100 tokens
Input price
$0.075 / 1M tokens
Output price
$0.25 / 1M tokens
Released
2026-08-27

Capabilities

  • Streaming
  • Function calling
  • Vision
  • JSON mode

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · discounts up to 50% · tier input average $0.075 / 1M tokens · tier output average $0.25 / 1M tokens

Z.ai

Tier: Standard · Region: SG · Quantization: fp8

Pricing
Input
$0.075 / 1M tokens
Output
$0.25 / 1M tokens
Cache read
$0.015 / 1M tokens
Discounts
  • Input: 50% off — list $0.15 / 1M tokens, discounted $0.075 / 1M tokens, effective $0.075 / 1M tokens
  • Cache read: 50% off — list $0.03 / 1M tokens, discounted $0.015 / 1M tokens, effective $0.015 / 1M tokens
  • Output: 50% off — list $0.5 / 1M tokens, discounted $0.25 / 1M tokens, effective $0.25 / 1M tokens
Data privacy and compliance
Region
SG
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
No
HIPAA compliant
No
SOC 2 certified
No
BYOK supported
No

Frequently asked questions

What is Z.AI: GLM 5.3 Flash?
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
How much does Z.AI: GLM 5.3 Flash cost?
Input costs start at $0.075 / 1M tokens and output costs start at $0.25 / 1M tokens. Provider-level prices vary by service tier.
Are provider discounts available for Z.AI: GLM 5.3 Flash?
Yes. Current provider offers include discounts of up to 50% from list price. The provider table shows list, discounted, and effective prices.
What is the context length of Z.AI: GLM 5.3 Flash?
Z.AI: GLM 5.3 Flash supports a 1,000,000 token context window and up to 131,100 output tokens.
What capabilities does Z.AI: GLM 5.3 Flash support?
Z.AI: GLM 5.3 Flash supports Streaming, Function calling, Vision, JSON mode.
Which providers offer Z.AI: GLM 5.3 Flash?
Z.AI: GLM 5.3 Flash is available from Z.ai.
How do providers handle data privacy for Z.AI: GLM 5.3 Flash?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models