Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

TL;DR AI
2 min readKey summary
Researchers introduced Gated DeltaNet-2, a linear-attention architecture with separate channel-wise erase and write gates for compressed memory updates.
By decoupling forgetting from writing, the model improves fast-weight editing while preserving efficient training and decoding.
In 1.3B-parameter experiments, it outperformed several competing architectures, including on long-context retrieval benchmarks like RULER.
The results suggest a practical fix for a core limitation of linear attention, with gains on language modeling and other NLP tasks.
