Moonshot: Kimi V1 128K

moonshotai/kimi-v1-128k

Kimi v1 (also marketed as the moonshot‑v1 family) is a large‑scale Transformer‑based language model developed by Moonshot AI that pushes the frontier of long‑context and multimodal AI by offering a maximum context window of 128 k tokens (roughly 200 k Chinese characters), automatically scaling to 8 k or 32 k modes for cost‑effective inference, and integrating visual understanding to process images directly; its core capabilities include extracting and summarizing massive documents in a single pass, generating structured outputs such as precise JSON through a dedicated JSON‑Mode, invoking external tools via ToolCalls, streaming partial results with Partial‑Mode, and performing real‑time internet searches to augment responses, while an auto‑context cache reduces token‑billing to just ¥1 per million cached tokens; technically, Kimi v1 leverages a ~2 trillion‑parameter Transformer architecture refined with reinforcement‑learning and long‑Chain‑of‑Thought (Long‑CoT) techniques, enabling high‑fidelity reasoning over ultra‑long texts, low‑latency token reuse, and robust multimodal perception, positioning it as one of the most advanced domestic LLMs for enterprise and developer applications.

Model specifications

Input
text
Output
text
Context
131,072 tokens
Max output
131,072 tokens
Input price
$2 / 1M tokens
Output price
$5 / 1M tokens
Released
2026-02-10

Capabilities

  • Streaming
  • Function calling
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $2 / 1M tokens · tier output average $5 / 1M tokens

Moonshot AI

Tier: Standard · Region: SG

Pricing
Input
$2 / 1M tokens
Output
$5 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
SG
Zero data retention
No
Data retention
Unknown retention
Used for training
Yes
Data collection
Moderated
No
GDPR compliant
No
HIPAA compliant
No
SOC 2 certified
No
BYOK supported
No

Frequently asked questions

What is Moonshot: Kimi V1 128K?
Kimi v1 (also marketed as the moonshot‑v1 family) is a large‑scale Transformer‑based language model developed by Moonshot AI that pushes the frontier of long‑context and multimodal AI by offering a maximum context window of 128 k tokens (roughly 200 k Chinese characters), automatically scaling to 8 k or 32 k modes for cost‑effective inference, and integrating visual understanding to process images directly; its core capabilities include extracting and summarizing massive documents in a single pass, generating structured outputs such as precise JSON through a dedicated JSON‑Mode, invoking external tools via ToolCalls, streaming partial results with Partial‑Mode, and performing real‑time internet searches to augment responses, while an auto‑context cache reduces token‑billing to just ¥1 per million cached tokens; technically, Kimi v1 leverages a ~2 trillion‑parameter Transformer architecture refined with reinforcement‑learning and long‑Chain‑of‑Thought (Long‑CoT) techniques, enabling high‑fidelity reasoning over ultra‑long texts, low‑latency token reuse, and robust multimodal perception, positioning it as one of the most advanced domestic LLMs for enterprise and developer applications.
How much does Moonshot: Kimi V1 128K cost?
Input costs start at $2 / 1M tokens and output costs start at $5 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Moonshot: Kimi V1 128K?
Moonshot: Kimi V1 128K supports a 131,072 token context window and up to 131,072 output tokens.
What capabilities does Moonshot: Kimi V1 128K support?
Moonshot: Kimi V1 128K supports Streaming, Function calling, JSON mode, Playground.
Which providers offer Moonshot: Kimi V1 128K?
Moonshot: Kimi V1 128K is available from Moonshot AI.
How do providers handle data privacy for Moonshot: Kimi V1 128K?
1 of 1 providers report that customer data is not used for training, and 0 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models