Tech News

Inkling shrinks to a quarter of its size

By Ratna Pratama
·
Share:
Inkling shrinks to a quarter of its size - inkling small
Inkling shrinks to a quarter of its size

Thinking Machines Lab has released a smaller version of its Inkling AI model that performs at nearly the same level while using far less computing power.

The new model, called Inkling-Small, contains 276 billion total parameters but only activates 12 billion at a time. That makes it roughly one-quarter the size of the original Inkling, which the company launched earlier this month following a partnership with Nvidia.

Performance matches larger model with less compute

According to the company, Inkling-Small matches or exceeds the original Inkling on reasoning and agentic tasks—areas where AI systems make decisions or take actions with minimal human input. The smaller model also processes 93 tokens per second, faster than the average for comparable models.

Independent benchmarks support these claims. Artificial Analysis gave Inkling-Small a 40% score on its Intelligence Index, placing it “well above average” among similar models. The original Inkling scored just one point higher at 41%, while DeepSeek V4 Flash tied Inkling-Small at 40%.

Related: Ireland braced for major solar eclipse

On more specialized tests, Inkling-Small outperformed its larger predecessor. It scored above 31% on Humanity’s Last Exam, compared to Inkling’s 29.7%. On SWEBench-Verified, a software engineering benchmark, it crossed 80%. The model accepts text, image, and audio inputs but outputs only text.

These results suggest that advances in model architecture can deliver similar performance without proportional increases in size or cost. For organizations running AI at scale, that could mean lower infrastructure demands without sacrificing capability.

Designed for real-world use cases

Thinking Machines said Inkling-Small is built for practical applications like image processing, document analysis, and data extraction from charts. The company emphasized that the model remains highly customizable, aligning with its broader mission to create AI that “extends human will and judgement.”

“Thinking Machines customers have seen first-hand that the right fine-tuned model can outperform closed models on a variety of tasks, and do so faster and cheaper,” the company said in a statement.

Related: Visa cuts 7% of workforce amid AI pivot

The original Inkling was trained across multiple domains—coding, reasoning, instruction-following, and audio-visual tasks—rather than being optimized for a single use case. Inkling-Small appears to maintain that versatility while improving efficiency.

Thinking Machines, led by Mira Murati, has positioned itself as a provider of flexible AI systems.

The partnership with Nvidia included access to GB300 NVL72 systems, which likely played a role in training both Inkling and Inkling-Small. The chipmaker also made a significant investment into the AI company.

While smaller models are often associated with trade-offs in accuracy or capability, Inkling-Small’s benchmarks suggest those gaps may be narrowing. For now, the company has not disclosed pricing or availability details for the new model.

Leave a Reply

Your email address will not be published. Required fields are marked *