
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks
Cursor Research open-sourced Mixture-of-Kittens (MoK), a deterministic MoE training megakernel for NVIDIA GB300 racks. The tool reportedly runs 2.37x faster than existing public baselines.
Cursor Research released MoK, an open-source megakernel designed for training Mixture-of-Experts models. It integrates communication and computation steps into a single deterministic process.
The kernel targets NVIDIA's GB300 NVL72 rack infrastructure. According to the release, it achieves speeds up to 2.37 times faster than the strongest available public baseline.
Optimizing MoE training is critical for scaling large language models efficiently. By reducing communication overhead, this tool could lower the computational cost of training complex architectures.
This release follows Cursor's work on Composer models. Open-sourcing the underlying training infrastructure allows other developers to leverage these optimizations for their own AI projects.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.