Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
TL;DR AI
2 min readKey summary
Researchers propose a geometry-based view of catastrophic forgetting in continually post-trained LLMs, arguing that update conflicts are relative to the model’s current state.
They validate the idea on Qwen3 models and introduce GCWM, which uses geometry conflict to gate Wasserstein merge corrections.
The method improves performance in continual learning settings and helps preserve prior capabilities without storing replay data.
