A new preprint on arXiv asks a pointed question: can out-of-distribution (OOD) generalization be predicted from a trained model's weights alone, without access to any target-domain data? The authors propose that circuit alignment—the degree to which internal computational pathways are structured consistently—might serve as such a predictor.

The paper positions this against existing representational similarity metrics like CKA, SVCCA, and RSA. Those methods compare activations across models or layers, but the new work appears to shift focus to the weights themselves, aiming for a forecast rather than a post-hoc comparison.

Because the abstract is brief, it does not include experimental results or conclusions. The contribution at this stage is the framing: a weight-based, target-free predictor of OOD behavior, distinct from activation-based similarity measures. Readers should treat the claims as preliminary until the full paper is available.