Z.AI: GLM 4.6V Flash

z-ai/glm-4.6v-flash

GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications. Details are as follows: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool use and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.

Model specifications

Input
text
Output
text
Context
200,000 tokens
Max output
128,000 tokens
Input price
$0.1 / 1M tokens
Output price
$0.1 / 1M tokens
Released
2026-02-17

Capabilities

  • Streaming
  • Function calling
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $0.1 / 1M tokens · tier output average $0.1 / 1M tokens

Z.ai

Tier: Standard · Region: SG

Pricing
Input
$0.1 / 1M tokens
Output
$0.1 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
SG
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
No
HIPAA compliant
No
SOC 2 certified
No
BYOK supported
No

Frequently asked questions

What is Z.AI: GLM 4.6V Flash?
GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications. Details are as follows: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool use and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.
How much does Z.AI: GLM 4.6V Flash cost?
Input costs start at $0.1 / 1M tokens and output costs start at $0.1 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Z.AI: GLM 4.6V Flash?
Z.AI: GLM 4.6V Flash supports a 200,000 token context window and up to 128,000 output tokens.
What capabilities does Z.AI: GLM 4.6V Flash support?
Z.AI: GLM 4.6V Flash supports Streaming, Function calling, JSON mode, Playground.
Which providers offer Z.AI: GLM 4.6V Flash?
Z.AI: GLM 4.6V Flash is available from Z.ai.
How do providers handle data privacy for Z.AI: GLM 4.6V Flash?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models