Switch language한국어
Back to the list

UCSD and Together AI Researchers Introduce Parcae: A Stable Architecture for Looped Language Models That Achieves the Quality of a Transformer Twice the Size

TL;DR AI

Key summary

2 min read
  1. UC San Diego and Together AI introduced Parcae, a middle-looped Transformer with built-in stability constraints.

  2. By treating the loop as a dynamical system and constraining the residual update, it avoids the training instability seen in earlier looped models.

  3. At tested scales, it outperforms matched baselines, including fixed-depth Transformers and prior recurrent depth models, with the same parameter budget.

  4. The approach can improve model quality without adding parameters or data, supporting cheaper inference and edge deployment.

Read the original