Switch language한국어
Back to the list

CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced CADENCE, a unified on-policy distillation framework for transferring reasoning into small language models.

  2. It combines adaptive KL scheduling with auxiliary mechanisms to reduce cold-start failures, improve scheduling, and address sparse rewards.

  3. On GSM8K and MATH-500, 0.5B students outperformed prior matched-compute baselines, using only a single Mac Studio.

  4. The work suggests a more stable, compute-efficient path for distilling math reasoning from large models into much smaller ones.

Read the original