Mira Murati's Thinking Machines Ships Inkling-Small: 276B Open-Weight Model Within One Point of Its Flagship


Thinking Machines Lab released Inkling-Small on July 30, two weeks after its 975B flagship Inkling, and the benchmark story is the interesting part: a model with fewer than one-third of the flagship's active parameters landing within a single point of it on the primary intelligence index.
The numbers: 276 billion total parameters, 12 billion active per token, using the same sparse Mixture-of-Experts architecture as the flagship. It achieves a score of 40 on the Artificial Analysis Intelligence Index, against the flagship's 41. On coding and frontier reasoning benchmarks — Humanity's Last Exam, GPQA Diamond, SciCode — it consistently meets or exceeds the larger model, trailing only slightly on agentic tasks and factual recall that benefit from scale.
The model is multimodal (text, image, and audio inputs), supports a 1 million-token context window, and is designed for organisations that want to fine-tune and deploy on their own infrastructure via the lab's Tinker platform. The licence is Apache 2.0 — commercially permissive, with no restriction on fine-tuning or redistribution.
Tinker API pricing for the flagship was reported at $1.87 per million input tokens and $4.68 per million output tokens (at a 64K context limit). Inkling-Small pricing has not been separately announced but is expected to come in below those figures given its inference efficiency advantage.
Thinking Machines Lab is less than a year old. Murati founded it after leaving OpenAI in late 2024 with a stated goal of building AI that organisations can inspect, customise, and control — rather than access through a closed API they don't control. The Inkling series is the first product expression of that thesis: open weights, a fine-tuning platform, and a flagship designed to be "broad and balanced" rather than a benchmark maximiser.
Inkling-Small fits the same philosophy at lower hardware cost. At 12B active parameters, it is self-hostable on significantly less VRAM than the full 41B-active flagship, making it the practical option for enterprises that want the Thinking Machines stack without the infrastructure overhead.


