Synthetic data has become a standard component of LLM training, especially for capabilities like autonomous and long-horizon task execution. But recent findings indicate that scaling up synthetic training data can hurt generation quality, according to the paper's abstract.
The new arXiv paper argues that the way synthetic data is curated matters. Specifically, it says effective curation requires group-level signals—information about collections of examples—rather than only individual example quality.
The abstract is brief and does not detail the proposed method or experiments. Still, it highlights a growing concern in the field: synthetic data is useful, but its benefits are not guaranteed at scale.