Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
TL;DR AI
2 min readKey summary
Researchers introduced Neuron-Activated Graph Ranking, a training-free method that ranks candidate pretraining data using sparse, high-impact neuron patterns from target examples.
Across six benchmarks, it outperformed random sampling and strong baselines, with notable gains on HellaSwag and in multi-target settings.
The analysis showed that the selected neurons were highly important to model performance, making the method more interpretable as well as effective.
This offers a more efficient way to choose target-relevant pretraining data for language models while improving benchmark results.
