Qwen: Qwen 3.5 Flash

qwen/qwen3.5-flash

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.

Model specifications

Input
text
Output
text
Context
1,000,000 tokens
Max output
64,000 tokens
Input price
$0.1 / 1M tokens
Output price
$0.4 / 1M tokens
Released
2026-02-26

Capabilities

  • Streaming
  • Function calling
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

2 available providers · tier input average $0.1 / 1M tokens · tier output average $0.4 / 1M tokens

Alibaba Cloud Int.(HK)

Tier: Standard · Region: HK

Pricing
Input
$0.1 / 1M tokens
Output
$0.4 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
HK
Zero data retention
No
Data retention
30-day retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
No

Alibaba Cloud Int.(SG)

Tier: Standard · Region: SG

Pricing
Input
$0.1 / 1M tokens
Output
$0.4 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
SG
Zero data retention
No
Data retention
30-day retention
Used for training
No
Data collection
Moderated
Yes
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
Yes

Frequently asked questions

What is Qwen: Qwen 3.5 Flash?
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.
How much does Qwen: Qwen 3.5 Flash cost?
Input costs start at $0.1 / 1M tokens and output costs start at $0.4 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Qwen: Qwen 3.5 Flash?
Qwen: Qwen 3.5 Flash supports a 1,000,000 token context window and up to 64,000 output tokens.
What capabilities does Qwen: Qwen 3.5 Flash support?
Qwen: Qwen 3.5 Flash supports Streaming, Function calling, JSON mode, Playground.
Which providers offer Qwen: Qwen 3.5 Flash?
Qwen: Qwen 3.5 Flash is available from Alibaba Cloud Int.(HK), Alibaba Cloud Int.(SG).
How do providers handle data privacy for Qwen: Qwen 3.5 Flash?
2 of 2 providers report that customer data is not used for training, and 0 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models