
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
Thinking Machines Lab released Inkling-Small, a 276B parameter multimodal MoE model with 12B active weights. The open-weights release matches predecessor performance at a quarter of the size and runs on one NVIDIA B300 GPU.
Thinking Machines Lab has unveiled Inkling-Small, an open-weights multimodal model utilizing a Mixture of Experts architecture. The system comprises 276 billion total parameters but activates only 12 billion during inference, optimizing computational efficiency.
According to the release, this smaller variant achieves performance levels comparable to the original Inkling model despite being significantly more compact. The reduction in size suggests a focus on balancing capability with resource consumption for broader usability.
A key technical detail involves the model's compatibility with NVIDIA hardware. The NVFP4 checkpoint allows the model to operate on a single B300 GPU, lowering the barrier for high-performance inference in enterprise or research settings.
By releasing open weights, the lab enables developers to fine-tune and deploy the architecture without relying on proprietary APIs. This move aligns with ongoing trends in the AI sector toward more efficient, accessible large language models.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.