Speculative decoding has become a popular way to speed up autoregressive language models: a small, fast drafter proposes candidate tokens, and the larger target model accepts or rejects them in batches. This preserves the target's output distribution while reducing the number of sequential decoding steps. The new arXiv paper, "Mentored Decoding," builds on this idea by explicitly allowing a controlled amount of drift from the target distribution, a technique known as lossy speculative decoding.

The authors describe their method as combining faster inference with boosting. Rather than treating the drafter as a mere approximation to be corrected, Mentored Decoding appears to use the drafter as a 'mentor' that guides the target model, with the lossy element enabling larger speed-ups. The abstract frames this as a way to push beyond the speed limits of exact speculative decoding while keeping the resulting quality degradation manageable.

Because the abstract is brief, the preprint does not yet provide experimental details or quantitative results in the visible text. Still, the conceptual link between boosting and speculative decoding is an interesting direction for the field, especially as researchers seek to balance latency, cost, and output fidelity in deployed LLMs.