Thinking Machines: Inkling
thinkingmachines/inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text.
Model specifications
- Input
- text, image, audio
- Output
- text
- Context
- 524,288 tokens
- Max output
- 524,288 tokens
- Input price
- $1 / 1M tokens
- Output price
- $4.05 / 1M tokens
- Released
- 2026-07-19
Capabilities
- Streaming
- Function calling
- Vision
- JSON mode
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
1 available provider · tier input average $1 / 1M tokens · tier output average $4.05 / 1M tokens
Together
Tier: Standard · Region: US
Pricing
- Input
- $1 / 1M tokens
- Output
- $4.05 / 1M tokens
- Cache read
- $0.17 / 1M tokens
No provider discount is currently published.
Data privacy and compliance
- Region
- US
- Zero data retention
- No
- Data retention
- Unknown retention
- Used for training
- No
- Data collection
- Moderated
- No
- GDPR compliant
- No
- HIPAA compliant
- No
- SOC 2 certified
- No
- BYOK supported
- No
Privacy policy · Terms · Official website · Documentation · Status · Support
Frequently asked questions
- What is Thinking Machines: Inkling?
- Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text.
- How much does Thinking Machines: Inkling cost?
- Input costs start at $1 / 1M tokens and output costs start at $4.05 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of Thinking Machines: Inkling?
- Thinking Machines: Inkling supports a 524,288 token context window and up to 524,288 output tokens.
- What capabilities does Thinking Machines: Inkling support?
- Thinking Machines: Inkling supports Streaming, Function calling, Vision, JSON mode.
- Which providers offer Thinking Machines: Inkling?
- Thinking Machines: Inkling is available from Together.
- How do providers handle data privacy for Thinking Machines: Inkling?
- 1 of 1 providers report that customer data is not used for training, and 0 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.