Modern LLM serving often routes a shared context across multiple models, such as when a coding agent switches models mid-session or a cascade escalates a request. Reusing the key-value (KV) cache from one model for another would save substantial compute, but caches are typically model-specific and do not transfer cleanly.

The new preprint, RaReCache, tackles this by selectively recomputing only the parts of the cache where the models' internal rankings diverge. Rather than discarding the entire cache or recomputing everything, the method uses rank disagreement as a signal for which tokens are likely to cause errors if reused.

The paper is an early-stage arXiv announcement, so details on the exact ranking metric and experimental results are not yet available in the abstract. Still, the approach points toward a practical middle ground for cross-model reuse, especially in settings where a shared context is passed between models with different architectures or training. As multi-model workflows become more common, such selective recomputation could help reduce latency without sacrificing output quality.