Thinking Machines drops 276B Inkling-Small with near-flagship scores and Apache 2.0 access
A quarter of Inkling’s size, fewer active parameters, and enterprise-friendly licensing, plus real pricing for 64K and 256K.

Thinking Machines, led by former OpenAI CTO Mira Murati, debuted Inkling-Small two weeks after launching its first open model, Inkling. The new 276B sparse multimodal model lands at 80.2% on SWE-bench Verified while shaving compute via 12B active parameters and offering Apache 2.0 licensing.
Two weeks after Thinking Machines released Inkling, its first open source AI language model, the company is back with Inkling-Small. The pitch is simple but high-stakes: a 276-billion-parameter multimodal reasoning model that gets within a single point of the original Inkling on a third-party Intelligence Index score, while using far less compute at inference time.
That “small” advantage is not marketing math. Inkling-Small has 12 billion active parameters per token. Inkling, its larger sibling, uses 41 billion active parameters per token. On the Artificial Analysis Intelligence Index, Inkling-Small scores 40, compared with 41 for Inkling. And on several targeted evaluations, it actually beats the bigger model: 80.2% on SWE-bench Verified versus 77.6% for Inkling, and 64.7% on Terminal Bench 2.1 versus 63.8%.
This matters because enterprises do not buy models the way gamers buy GPUs. They buy compute budgets, deployment predictability, and legal clarity. Thinking Machines is explicit about the operational trade: the model is still far too large for a laptop or conventional workstation, but it is materially easier to run than the 3.5X larger flagship because the active parameter count drops dramatically. The practical angle is that organizations with some, but not a lot, of their own GPUs could consider self-hosting more realistically than with the full flagship. In a world where hosting costs and latency are the true “hidden taxes” of AI adoption, smaller active-parameter footprints can be the difference between pilots and production.
Now zoom in on the economics Thinking Machines is advertising at launch. The company released full weights on Hugging Face and added support for fine-tuning via its Tinker model training API. It also offered a limited-time 50% discount. For the standard 64K-context Inkling-Small model, API pricing is listed at $0.58 per million prefill (input) tokens, $1.44 per million sampled (output) tokens, and $1.73 per million training tokens. Cached prefill requests are priced at $0.116 per million tokens. A 256K-context variant exists too, at higher rates.
The “near-flagship performance” story gets even more interesting when you look at what Inkling-Small is architecturally doing. It is a sparse Mixture-of-Experts model. Thinking Machines says it uses a 42-layer decoder that routes each token to six of 256 specialized experts, plus two shared experts that remain active for every token. That design helps explain the big gap between total parameters (276B) and active parameters (12B). In other words, the model carries a large pool of learned capacity, but it activates only part of it per inference step. It also is natively multimodal: images, audio, and text are projected into a shared representation and processed jointly by the decoder, rather than being managed through fully separate external systems. The company lists coding assistants, agentic applications, chatbots, RAG systems, and other multimodal applications as intended uses.
But let’s not pretend “small” equals easy. Thinking Machines says the standard BF16 checkpoint needs at least 600 GB of aggregate GPU memory. It lists supported configurations as 4x NVIDIA B300 GPUs or 8x NVIDIA H200 GPUs. A quantized NVFP4 checkpoint lowers the requirement to roughly 180 GB of aggregate VRAM. The company says the NVFP4 version can run in W4A4 mode on a single NVIDIA B300, or in W4A16 mode on two H200 GPUs. That still rules out ordinary laptops, MacBooks, desktop gaming PCs, and most developer workstations. So the target is enterprise GPU servers, cloud clusters, and specialized inference providers. The relative win is still real: less memory translate into lower hosting costs, easier capacity planning, and a wider group of organizations able to self-host.
The enterprise differentiator may be licensing even more than benchmarks. Inkling-Small is released under Apache 2.0, a permissive license that generally allows organizations to use, modify, fine-tune, redistribute, and commercialize the model, including inside proprietary products, subject to notice and attribution requirements. That kind of clarity can simplify adoption for legal, procurement, and platform teams, compared with custom “open” AI licenses that sometimes come with extra commercial conditions or restrictions. Thinking Machines’ choice does not remove the need for acceptable-use policies, data provenance checks, regulatory exposure review, or downstream safety obligations, but Apache 2.0 provides a cleaner starting point for building internal systems, shipping commercial products, and maintaining modified versions.
Finally, the performance story has a nuance board members and product leaders will care about. Inkling-Small is not a drop-in replacement for every workload. The gains are not universal. Inkling retains a clear advantage on factual knowledge and some agentic tasks. Inkling-Small scores 15.5% on τ³-Banking versus 23.7% for Inkling, and its AA Omniscience score is negative, reflecting weaker factual coverage. The reported hallucination rate is slightly lower for Inkling-Small, but factual coverage is still the trade. Thinking Machines’ enterprise implication is straightforward: for high-stakes factual tasks, organizations using Inkling-Small will still need retrieval, verification, and human review.
Taken together, Inkling-Small is a signal that open model releases are moving into a more operationally disciplined phase: repeatable engineering, sparse compute, multimodal natively supported inputs, and licenses that enterprises can actually work with. For leaders evaluating build-versus-buy and model refresh cycles, this is a reminder that “best model” is increasingly “best deployable model,” and that shaving active compute while keeping benchmark strength is where competitive advantage hides.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology
Larry Ellison is sprinting on debt to make Oracle the AI face
Oracle's founder is pushing a debt-fueled rebuild of his data empire for the AI boom. For leaders, it signals risk and leverage tradeoffs.

Falcon 9 will deliberately crash into the Moon, and astronomers can likely spot it
SpaceX’s planned lunar impact is expected to loft a high debris plume visible through some telescopes, drawing live scientific attention.

Starship’s Flight 13 deployed 20 Starlink V3 satellites, captured 65-second heat-shield video
A new Starlink V3 camera view shows Ship firing Raptors as 100,000-satellite ambitions draw closer.

