LLM Serving Fairness: No More Noisy Neighbors | Cohere

TL;DR AI
1 min readKey summary
Cohere outlined a fairness system for shared LLM inference on GPUs.
The design combines admission control, service tiers, and Deficit Round Robin scheduling.
It is meant to stop one tenant’s traffic spikes from crowding out others.
The system preserves batching efficiency while keeping priority and deadline ordering.

