Switch language한국어
Back to the list

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

TL;DR AI

Key summary

2 min read
  1. Researchers built a controlled multi-turn environment to study long-horizon planning across three stages: pre-training, single-teacher post-training, and multi-teacher integration.

  2. Explicit world-model construction and some long-horizon data helped agents learn planning, while suboptimal trajectories hurt performance.

  3. In post-training, OPD had a wider useful range than GRPO in difficult settings, and distilling incompatible teacher knowledge could degrade prior understanding.

  4. Multi-teacher on-policy distillation worked best when environments shared compatible planning patterns; conflicting patterns led to interference.

Read the original