Tuesday, 22 September 2026

Search
Latent Digest

TECHNOLOGY, TRACKED ACROSS DISCIPLINES

Research Digest

New Studies Probe How LLMs Store, Forget, and Distort Knowledge

Three papers examine the mechanics of knowledge in large language models, from tracing its origins to removing it and warning against its misuse.

· 1 min read · 3 sources

Three recent papers converge on a central theme: knowledge in large language models is not a simple lookup table but a complex, manipulable property of their internal representations. The first, LMEnt, introduces a suite for analyzing how knowledge flows from pretraining data into model representations, providing tools to trace the origins of what a model knows. The second, KUDA, takes a more interventionist approach, proposing a method to unlearn specific knowledge by deviating the model's representation of that knowledge—essentially editing what the model believes.

Where these two agree is that knowledge is encoded in representations and can be systematically studied or altered. They differ in purpose: LMEnt is diagnostic, while KUDA is corrective. A third study, accepted for publication in AI and Ethics, adds a cautionary note. It empirically shows that LLMs can erode scientific understanding, potentially misaligning users' grasp of science. This suggests that the same representational mechanisms that enable knowledge storage and unlearning can also produce harmful distortions.

Taken together, the papers highlight a growing maturity in the field: researchers are moving beyond simply training models to understanding and controlling what they know—and guarding against what they get wrong. The contrast between the constructive tools (LMEnt, KUDA) and the risk-focused study underscores the dual-use nature of knowledge manipulation in LLMs.

Sources · 3

  1. 01LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to RepresentationsarXiv
  2. 02KUDA: Knowledge Unlearning by Deviating Representation for Large Language ModelsarXiv
  3. 03Large language models eroding science understanding: an empirical study of malignmentarXiv

More in Research Digest