Knowledge editing is usually judged by three tests: did the targeted fact change, do paraphrases reflect the edit, and are unrelated answers left intact. These checks are meant to confirm that an edit took hold, generalizes, and does not cause collateral damage.
A new arXiv preprint argues that a model can pass all three and still lose something none of them measures: the ability to tell good evidence from bad. The paper specifically points to sequential knowledge editing, where multiple edits are applied in order, as a process that can break this ability without showing up in standard accuracy metrics.
The implication is that current evaluation protocols have a blind spot. A model may appear unchanged by conventional measures while its reasoning about evidence quality has silently degraded, leaving a gap between what benchmarks certify and what the model can actually do.