Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
TL;DR AI
2 min readKey summary
Researchers introduced Word Coverage Score (WCS) to measure how sampling filters in LLM decoding can block otherwise valid words.
Using open-weight models, they showed that common settings like top-p, top-k, and min-p can prune contextually appropriate vocabulary.
The findings suggest decoding choices may trade off coherence for lexical richness, reducing expressive diversity in generated text.
