New research posted on arXiv examines why sparse autoencoders (SAEs) sometimes produce fragmented representations of visual concepts. SAEs are trained to decompose model activations into sparse combinations of interpretable dictionary atoms, an approach grounded in the linear representation hypothesis (LRH). The paper argues that the SAE objective goes beyond LRH by smuggling in an additional assumption: an independence prior over the learned atoms.

That prior, the authors contend, is what fragments visual concepts. Rather than encoding a concept as a single coherent direction, the SAE may split it across multiple atoms that are individually sparse and independent but do not correspond to the underlying visual structure. The result is a dictionary that is interpretable in isolation yet fails to reflect how concepts actually compose in the model.

The paper is a single preprint, and the argument is presented as a critique of the standard SAE formulation rather than a new training method. The authors do not claim that all SAEs behave this way, but they suggest the independence prior is baked into the objective and should be examined when interpreting SAE results.