AI
[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B
Thinking Machines Lab has launched Inkling, its first fully released open-weights foundation model family.
Key takeaways
- Inkling is a Mixture-of-Experts model with 975B total parameters and 41B active parameters.
- Inkling supports a context window of up to 1M tokens and was pretrained on 45 trillion tokens.
- Inkling is the leading U.S. open-weights release, scoring 41 on the Intelligence Index.
- Inkling is a clear step up from Nemotron Ultra but remains behind GLM 5.2 on agentic benchmarks.
- Inkling-Small is unexpectedly competitive versus the larger model on several evaluations.
Thinking Machines Lab has launched Inkling, its first fully released open-weights foundation model family. Inkling is a multimodal Mixture-of-Experts transformer with 975B total parameters (41B active) and supports up to a 1M token context window. It was pretrained on 45 trillion tokens of text, images, audio, and video under an Apache 2.0 license. Alongside it, the lab shared a preview of Inkling-Small, a lighter-weight model with 12B active parameters. Independent commentators and benchmarks, such as Artificial Analysis, have positioned Inkling as the strongest U.S.-based open-weight release to date, though it still trails some top Chinese open-weight models on specific benchmarks.
By the numbers
- 975B
- Total parameters in the Inkling model
- 41B
- Active parameters in the Inkling model
- 45T
- Tokens of text, images, audio, and video pretrained on
- 1M
- Maximum context window supported by open-weights checkpoints
- 276B
- Total parameters in the Inkling-Small model
How it unfolded
- Thinking Machines Lab launches Inkling and previews Inkling-Small
Turn stories like this into views
Ravenclip finds the AI news, makes the video, and posts it before attention moves on.
Common questions
- What happened with Inkling?
- Inkling is a Mixture-of-Experts model with 975B total parameters and 41B active parameters.
- Where can I read the original report?
- Read the full report at latent_space.