Meituan: LongCat Flash Chat

meituan/longcat-flash-chat

LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce communication overhead and achieve high throughput while maintaining training stability through advanced scaling strategies such as hyperparameter transfer, deterministic computation, and multi-stage optimization. This release, LongCat-Flash-Chat, is a non-thinking foundation model optimized for conversational and agentic tasks. It supports long context windows up to 128K tokens and shows competitive performance across reasoning, coding, instruction following, and domain benchmarks, with particular strengths in tool use and complex multi-step interactions.

Model specifications

Input
text
Output
text
Context
131,072 tokens
Max output
32,800 tokens
Input price
$0.2 / 1M tokens
Output price
$0.8 / 1M tokens
Released
2026-03-01

Capabilities

  • Streaming
  • Function calling
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $0.2 / 1M tokens · tier output average $0.8 / 1M tokens

AtlasCloud

Tier: Standard · Region: US

Pricing
Input
$0.2 / 1M tokens
Output
$0.8 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
No
Data retention
7-day retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
No
HIPAA compliant
No
SOC 2 certified
No
BYOK supported
No

Frequently asked questions

What is Meituan: LongCat Flash Chat?
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce communication overhead and achieve high throughput while maintaining training stability through advanced scaling strategies such as hyperparameter transfer, deterministic computation, and multi-stage optimization. This release, LongCat-Flash-Chat, is a non-thinking foundation model optimized for conversational and agentic tasks. It supports long context windows up to 128K tokens and shows competitive performance across reasoning, coding, instruction following, and domain benchmarks, with particular strengths in tool use and complex multi-step interactions.
How much does Meituan: LongCat Flash Chat cost?
Input costs start at $0.2 / 1M tokens and output costs start at $0.8 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Meituan: LongCat Flash Chat?
Meituan: LongCat Flash Chat supports a 131,072 token context window and up to 32,800 output tokens.
What capabilities does Meituan: LongCat Flash Chat support?
Meituan: LongCat Flash Chat supports Streaming, Function calling, JSON mode, Playground.
Which providers offer Meituan: LongCat Flash Chat?
Meituan: LongCat Flash Chat is available from AtlasCloud.
How do providers handle data privacy for Meituan: LongCat Flash Chat?
1 of 1 providers report that customer data is not used for training, and 0 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models