Z.AI: GLM 4.5 Flash
z-ai/glm-4.5-flash
GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters. Both models share a similar training pipeline: an initial pretraining phase on 15 trillion tokens of general-domain data, followed by targeted fine-tuning on datasets covering code, reasoning, and agent-specific tasks. The context length has been extended to 128k tokens, and reinforcement learning was applied to further enhance reasoning, coding, and agent performance. GLM-4.5 and GLM-4.5-Air are optimized for tool invocation, web browsing, software engineering, and front-end development. They can be integrated into code-centric agents such as Claude Code and Roo Code, and also support arbitrary agent applications through tool invocation APIs. Both models support hybrid reasoning modes, offering two execution modes: Thinking Mode for complex reasoning and tool usage, and Non-Thinking Mode for instant responses. These modes can be toggled via the thinking.typeparameter (with enabled and disabled settings), and dynamic thinking is enabled by default.
Model specifications
- Input
- text
- Output
- text
- Context
- 128,000 tokens
- Max output
- 96,000 tokens
- Input price
- $0.1 / 1M tokens
- Output price
- $0.1 / 1M tokens
- Released
- 2026-02-17
Capabilities
- Streaming
- Playground
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
1 available provider · tier input average $0.1 / 1M tokens · tier output average $0.1 / 1M tokens
Z.ai
Tier: Standard · Region: SG
Pricing
- Input
- $0.1 / 1M tokens
- Output
- $0.1 / 1M tokens
No provider discount is currently published.
Data privacy and compliance
- Region
- SG
- Zero data retention
- Yes
- Data retention
- Zero retention
- Used for training
- No
- Data collection
- Moderated
- No
- GDPR compliant
- No
- HIPAA compliant
- No
- SOC 2 certified
- No
- BYOK supported
- No
Privacy policy · Terms · Official website · Documentation · Support
Frequently asked questions
- What is Z.AI: GLM 4.5 Flash?
- GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters. Both models share a similar training pipeline: an initial pretraining phase on 15 trillion tokens of general-domain data, followed by targeted fine-tuning on datasets covering code, reasoning, and agent-specific tasks. The context length has been extended to 128k tokens, and reinforcement learning was applied to further enhance reasoning, coding, and agent performance. GLM-4.5 and GLM-4.5-Air are optimized for tool invocation, web browsing, software engineering, and front-end development. They can be integrated into code-centric agents such as Claude Code and Roo Code, and also support arbitrary agent applications through tool invocation APIs. Both models support hybrid reasoning modes, offering two execution modes: Thinking Mode for complex reasoning and tool usage, and Non-Thinking Mode for instant responses. These modes can be toggled via the thinking.typeparameter (with enabled and disabled settings), and dynamic thinking is enabled by default.
- How much does Z.AI: GLM 4.5 Flash cost?
- Input costs start at $0.1 / 1M tokens and output costs start at $0.1 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of Z.AI: GLM 4.5 Flash?
- Z.AI: GLM 4.5 Flash supports a 128,000 token context window and up to 96,000 output tokens.
- What capabilities does Z.AI: GLM 4.5 Flash support?
- Z.AI: GLM 4.5 Flash supports Streaming, Playground.
- Which providers offer Z.AI: GLM 4.5 Flash?
- Z.AI: GLM 4.5 Flash is available from Z.ai.
- How do providers handle data privacy for Z.AI: GLM 4.5 Flash?
- 1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.