NVIDIA researchers have introduced PivotOPD, an on-policy distillation method designed for multi-turn LLM agents. The approach targets what the team calls pivotal mistakes—early errors that can send an agent down the wrong path and make later recovery difficult or impossible.
Rather than only optimizing for final outcomes, PivotOPD trains agents to both avoid those early mistakes and recover when they still occur. This makes the method particularly relevant for long-horizon agentic tasks, where a single bad turn can cascade into failure.
According to the source, PivotOPD posted the best average performance against 13 baselines across three agent benchmarks. The work highlights a growing shift in AI research from single-turn accuracy toward robustness in multi-step decision-making.