AI Models & Platforms

Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model

mm
Add Unite.AI to your preferred sources on Google

Cohere released North Small Translate on September 10, 2026, an open-weight mixture-of-experts machine translation model with 218 billion total parameters and 25 billion active parameters, licensed for research and non-commercial use under CC BY-NC 4.0.

Cohere described North Small Translate as its first translation model in the North model family, building on a multilingual and translation lineage running from the Tiny Aya model family through Command A Translate. The model is text-only in both directions, with a 16K-token input context and a 16K-token output limit, and Cohere describes support for more than 50 languages. Minimum hardware is one B200 GPU or two H100 GPUs at W4A4 quantization.

Cohere-Reported Benchmarks, Throughput, and Cost

Cohere reported a WMT26 all-languages score of 83.60 for North Small Translate, rising to 84.36 for an agentic multi-pass variant that the company says can find and fix errors in translation. In Cohere’s evaluation, which used GPT-5.6-Sol as a judge, Qwen 3.5 397B A17B scored 81.56, DeepL NextGen 81.37, Gemma 4 31B (on) 79.46, GLM 5.2 FP8 76.50, and Google Translate 68.20.

The WMT benchmark evaluates the general capabilities of machine translation systems across a wide range of languages, domains, genres, and modalities, and its scoring methodology treats scores from 80 to 100 as perfect or carrying only minor errors. Cohere described North Small Translate as the best-performing dedicated machine translation model in the evaluation, open or closed, on average across all languages, and said it performs strongly across 32 high-resource languages and 18 additional languages.

At the regional level, Cohere reported the model beating Gemma 4 31B (on) in Europe, 82.2 to 73.9, while running essentially even with it in South Asia, 86.2 to 86.7. Cohere said both versions also outperformed DeepL NextGen in every non-European region tested: by roughly 8 to 10 points in South Asia and MENA, about 4 to 5 points in Southeast Asia, and 1 to 3 points in East Asia, where the standard model approached DeepL’s 85.41 score.

On Cohere’s long-context evaluation, which measures how well a model translates two chapters of a book in a single call with per-paragraph quality scored through xComet-XL, North Small Translate scored 48.9 against 21.3 for Google Translate and 19.4 for Gemma 4 31B. Cohere described that result as more than double both baselines and ahead of every general-purpose LLM it tested.

In internal throughput testing under identical concurrency levels and hardware configurations, Cohere measured up to 1.4 times higher output throughput than Gemma 4 31B TP1: 112 versus 81 output tokens per second at low concurrency and 39 versus 30 at high concurrency, which the company put at 30 to 38 percent more tokens generated per second.

For enterprises evaluating commercial licenses, Cohere reported an 80.1 score at $0.000676 per task using 661 tokens on average. The company listed Gemini 3.1 Pro Preview (high) at $0.038928 per task, stated as 5,762 percent more, and Qwen 3.5 397B A17B and its own Command A+ at $0.004525 and $0.005158 per task, using per-token or per-character pricing from model providers’ websites.

Architecture and Language Coverage

The model card describes North Small Translate as a decoder-only sparse mixture-of-experts Transformer with 128 experts, of which eight activate per token, alongside shared experts applied to every token. Attention layers interleave sliding-window attention (window size 4096, using Rotary Positional Embeddings) with global attention layers without positional embeddings in a 3:1 ratio, a design first introduced in Command A. The router applies a sigmoid activation over the expert logits and normalizes over the selected top-k. The model was post-trained specifically for translation quality.

Where the announcement describes support for more than 50 languages, the model card enumerates exactly 50, including English, Arabic, French, German, Hindi, Japanese, Korean, Russian, Ukrainian, Vietnamese, and Simplified and Traditional Chinese. Cohere’s documentation describes the coverage as English plus more than 50 language and locale variants, names Modern Standard Arabic, German, French, Japanese, Korean, Russian, and Ukrainian as tier-one languages, and lists locale variants such as Egyptian, Modern Standard, and Saudi Arabic, Portugal Portuguese, Cyrillic Serbian, and Norwegian Bokmål.

Licensing, Access, and the RWS Partnership

The weights are gated on Hugging Face: downloaders must agree to share their contact information, and use is governed by the CC BY-NC 4.0 license together with Cohere Labs’ Acceptable Use Policy. Three quantization variants are available: BF16 (four B200 or eight H100 GPUs), FP8 (two B200 or four H100), and NVFP4 W4A16 (one B200 or two H100), which the model card states are the checkpoints Cohere serves in production. A hosted Hugging Face Space offers a demo of the model.

The model is also served on Cohere’s API free tier through the Chat V2 API under the model ID north-small-translate-1-0, free until rate limits are reached, with commercial licensing and deployment available through Cohere’s Model Vault at suggested hardware of two H100 GPUs or one B200. Cohere’s documentation lists use cases including knowledge management, safety and operations documentation, internal communication, and localization and customer support.

Cohere said North Small Translate was developed in partnership with RWS, an AI solutions company in language technology and services whose Language Weaver research and science teams and language experts helped shape the model’s real-world translation performance throughout development. The company states that RWS works with more than 80 percent of the world’s top 100 brands, and that enterprises needing a dedicated translation and localization platform can access the model through RWS’s Language Weaver product.

According to Cohere’s June 1, 2026 account of the collaboration, Cohere asked RWS to test Command A Translate in July 2025, and by September 2025 the two teams had begun work on a specialized translation model that the post said came to power RWS’s Language Weaver Pro. Cohere reported that Language Weaver Pro achieved 55 percent overall sentence-level wins against DeepL NextGen in human evaluations and outperformed competitors in automated benchmarks in 31 of 32 languages commonly used across enterprise domains. RWS CEO Ben Faes described the model behind Language Weaver Pro as the brain behind the product.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.