PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Abstract
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Community
TL;DR: Multi-turn agents don't fail everywhere. Most failed rollouts hinge on one early pivotal mistake, the mistake is usually recoverable, and standard on-policy distillation can't repair it because the student never samples the recovery action. PivotOPD fixes that by concentrating distillation at the pivotal turn and the turns right after it.
What we found (ALFWorld, oracle replay of failed Qwen3 8B–235B rollouts)
- 59% of failed rollouts contain a pivotal mistake, an action that moves the agent farther from the goal, at a median of turn 8–12 out of 30. The rest of the episode is wasted.
- Correcting that one turn lifts replayed success from 8% to 59%. Leaving the mistake in place and guiding only the two turns after it still reaches 58%.
- Standard OPD cuts overall failures from 79% to 56%, but failures after a pivotal mistake barely move (51% → 49%). The recovery action stays below 1% probability, so a group of 8 rollouts never samples it: no sample, no learning signal.
- Pivot detection. A teacher model reads each rollout in hindsight, picks candidate turns and names a gold action at each; a turn is pivotal when the student's action disagrees. The teacher then names a recovery action for each of the next K turns. It only ever names actions.
- Preventive distillation (reverse KL). The student's own response at the pivotal turn is re-scored against a privileged self-teacher, the frozen student hinted with the gold action.
- Recovery distillation (forward KL). The self-teacher writes recovery responses from the post-mistake state, and the unhinted student is trained on them. Forward KL is mass-covering, so it places probability on actions the student would never sample on its own.
Both terms enter a single PPO update alongside standard group-based RL.
Results
- Against 13 baselines on ALFWorld, WebShop and Search-based QA, PivotOPD wins all 8 per-benchmark averages for both Qwen3-1.7B and Qwen3-8B students (+5.5% / +5.9% over the strongest baseline on ALFWorld / Search-based QA with the 1.7B student).
- Replayed from the same 72 pivotal mistakes, it recovers 72.7% of the time vs. 8.3% for the base model and 20.3% for standard OPD, and in the fewest turns.
- It doesn't need a bigger teacher: with Qwen3-8B as its own teacher it's still best on all three benchmarks (+3.9% on average).
- The same principle extends to software engineering: a Nemotron-3.5 student with a Nemotron-3-Super teacher gains +3.2 points on SWE-Bench Verified, where standard OPD gains +0.2.
🌐 Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Get this paper in your agent:
hf papers read 2609.40285 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
