Nous Research Releases Token Superposition Training to Speed Up LLM Pre-Training by Up to 2.5x Across 270M to 10B Parameter Models

TL;DR AI
2 min readKey summary
Nous Research introduced Token Superposition Training, a two-stage pre-training method that compresses early learning by grouping tokens into bags before returning to standard next-token prediction.
Across 270M to 10B parameter models, TST matched or beat equal-compute baselines while reducing total training time.
The biggest reported gain was about 2.5x faster training on a 10B-A1B mixture-of-experts run.
If these results generalize, TST could significantly lower the cost and wall-clock time of training large language models without changing final behavior.
