A new blog post from Google Research contends that the biggest limitation on large language models' factual accuracy is not how much they know, but how well they can pull up the right knowledge at the right time. The authors draw an analogy between empty shelves and lost keys: a model may have plenty of storage space (shelves) but still fail because it cannot locate the specific fact (lost keys) when asked.

This distinction matters because it shifts the focus from scaling model size to refining how models retrieve information. The post suggests that even when a fact is encoded in the parameters, the model's inference-time process often fails to surface it. This explains why a model might answer a simple factual question incorrectly despite having been trained on the answer.

The researchers propose that improving recall—through better prompting, decoding strategies, or attention mechanisms—could be a more effective path to factual reliability than simply adding more parameters. This challenges the assumption that bigger models automatically become more truthful, pointing instead to the need for targeted work on how knowledge is accessed during generation.