Z.AI: GLM 5.3 Flash
z-ai/glm-5.3-flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
Model specifications
- Input
- text, image, video
- Output
- text
- Context
- 1,000,000 tokens
- Max output
- 131,100 tokens
- Input price
- $0.075 / 1M tokens
- Output price
- $0.25 / 1M tokens
- Released
- 2026-08-27
Capabilities
- Streaming
- Function calling
- Vision
- JSON mode
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
1 available provider · discounts up to 50% · tier input average $0.075 / 1M tokens · tier output average $0.25 / 1M tokens
Z.ai
Tier: Standard · Region: SG · Quantization: fp8
Pricing
- Input
- $0.075 / 1M tokens
- Output
- $0.25 / 1M tokens
- Cache read
- $0.015 / 1M tokens
Discounts
- Input: 50% off — list $0.15 / 1M tokens, discounted $0.075 / 1M tokens, effective $0.075 / 1M tokens
- Cache read: 50% off — list $0.03 / 1M tokens, discounted $0.015 / 1M tokens, effective $0.015 / 1M tokens
- Output: 50% off — list $0.5 / 1M tokens, discounted $0.25 / 1M tokens, effective $0.25 / 1M tokens
Data privacy and compliance
- Region
- SG
- Zero data retention
- Yes
- Data retention
- Zero retention
- Used for training
- No
- Data collection
- Moderated
- No
- GDPR compliant
- No
- HIPAA compliant
- No
- SOC 2 certified
- No
- BYOK supported
- No
Privacy policy · Terms · Official website · Documentation · Support
Frequently asked questions
- What is Z.AI: GLM 5.3 Flash?
- GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.
- How much does Z.AI: GLM 5.3 Flash cost?
- Input costs start at $0.075 / 1M tokens and output costs start at $0.25 / 1M tokens. Provider-level prices vary by service tier.
- Are provider discounts available for Z.AI: GLM 5.3 Flash?
- Yes. Current provider offers include discounts of up to 50% from list price. The provider table shows list, discounted, and effective prices.
- What is the context length of Z.AI: GLM 5.3 Flash?
- Z.AI: GLM 5.3 Flash supports a 1,000,000 token context window and up to 131,100 output tokens.
- What capabilities does Z.AI: GLM 5.3 Flash support?
- Z.AI: GLM 5.3 Flash supports Streaming, Function calling, Vision, JSON mode.
- Which providers offer Z.AI: GLM 5.3 Flash?
- Z.AI: GLM 5.3 Flash is available from Z.ai.
- How do providers handle data privacy for Z.AI: GLM 5.3 Flash?
- 1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.