A new preprint on arXiv examines why language models remain coherent when their internal activations are perturbed. The authors point to interventions like linear steering, which can shift activations without fully derailing generated text, and ask what underlying property explains this resilience.

The paper's hypothesis is that robustness comes from "privileged error-correcting basins": regions of activation space that naturally pull perturbed states back toward coherent behavior. The abstract frames this as a passive property, not an active mechanism the model learns to defend itself.

Because this is a single preprint, there are no differing claims to reconcile. The proposed basin-based explanation is presented as a hypothesis, and the full evidence will be in the paper's body.