Data selection has become a central bottleneck in large-language-model training: web-scale corpora are noisy, and token budgets are limited. A new arXiv paper argues that the problem is especially hard in continual pre-training (CPT), where poor data choices can cause the model to lose knowledge it already has. In this setting, data selection becomes a forgetting-control problem rather than simply a matter of picking the most informative examples.
The paper's proposed approach combines Fisher information with submodular optimization. As the title suggests, the method is model-aware: it uses a sensitivity signal from the model itself to guide which data are selected, while submodularity makes the subset-selection problem more tractable and encourages useful coverage of the data.
The preprint is framed around the CPT bottleneck. Because the source is an abstract, there is little public detail on benchmark results or comparisons, but the motivation is clear: better data selection, guided by the model's own sensitivity, may reduce forgetting when a pretrained model is trained further.