Cohere's open-weight ASR model hits 5.4% word error rate — low enough to replace speech APIs in production pipelines

TL;DR AI
2 min readKey summary
Cohere's Transcribe model has 2 billion parameters and is released under Apache-2.0.
It achieves an average word error rate (WER) of 5.42%.
Trained on 14 languages, including English, Chinese, Japanese, Korean and Arabic.
It outperforms Whisper and ElevenLabs and currently tops the Hugging Face ASR leaderboard.
Available for commercial use, accessible via API or in Cohere’s Model Vault as cohere-transcribe-03-2026, and can run on an organization’s own infrastructure.
