Retrieval-based speculative decoding speeds up language model generation by drafting tokens from text that already exists, rather than predicting each token from scratch. This approach is particularly relevant for coding agents, which repeatedly generate similar code, logs, and earlier attempts. A new preprint on arXiv, titled AgSpec, focuses on exactly this setting.

The abstract notes that existing retrieval-based methods fall short in agentic pipelines. AgSpec is presented as a way to push the limits of this technique, though the visible abstract does not include specific benchmark numbers or implementation details.

Because only one source is available, there are no conflicting accounts to compare. The claims here are limited to what the paper itself describes.