Cohere launched North Mini Code on September 5, introducing an open-source mixture-of-experts (MoE) coding model with 30 billion total parameters and 3 billion active parameters designed for software development and agentic coding tasks [1, 2]. The model delivers strong performance while maintaining efficiency that allows it to run on practical hardware setups, including a single Nvidia H100 GPU, enabling developers to self-host without needing multi-GPU infrastructure [2].
North Mini Code scored 33.4 on the Artificial Analysis Coding Index, outperforming other open-weight models of similar size such as Alibaba’s Qwen3 and Google’s Gemma 4 [1, 2]. In testing, it achieved up to 2.8 times higher output throughput and 30% lower inter-token latency compared to Devstral Small 2 when run under the same conditions [1].
The model is freely available under the Apache 2.0 license and can be accessed through direct download, a managed inference environment, or API [1, 2]. Nick Frosst, Cohere co-founder, emphasized the company's focus on sovereign AI, saying, "We’re now hearing similar concerns from developers. They’re starting to think of model access as infrastructure, and infrastructure should be something you own and control. That is an extension of sovereignty." He added, "We want to give developers a capable, fast, open-weight model they can run locally on their own terms, and that fits in their compute environments." [2]
Cohere’s release of North Mini Code targets developers seeking flexible, performant code generation tools that do not require large-scale centralized compute resources. Access under a permissive license aims to ensure control and ownership of these AI tools remain with users.
North Mini Code’s availability today marks the start of accessible agentic coding models that combine scale and efficiency, designed for integration in local and cloud-based developer workflows [1, 2].