HRM-Text: Efficient Pretraining Beyond Scaling
TL;DR AI
2 min readKey summary
Researchers introduced HRM-Text, a 1B-parameter hierarchical recurrent language model trained from scratch on 40 billion unique tokens.
It uses instruction-response pairs plus stabilization methods like MagicNorm and PrefixLM masking to make deep recurrence train reliably.
HRM-Text achieved competitive results on benchmarks including MMLU, ARC-C, DROP, GSM8K, and MATH.
The model appears to use far fewer tokens and much less estimated compute than typical Transformer-based pretraining baselines.
The work suggests architecture and objective changes can dramatically lower the cost of training strong language models.
