Production-Ready W4A8: vLLM Integration Explained | Cohere

TL;DR AI
2 min readKey summary
Cohere unveiled production-ready W4A8 inference kernels for NVIDIA Hopper GPUs.
The kernels are integrated with vLLM and support both dense and MoE models.
A LUT-based dequantization approach helps overcome FP8 and INT4 conversion limits.
Compared with W4A16, the system reportedly cuts time to first token by up to 58% and time per output token by up to 45%.

