RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
TL;DR AI
2 min readKey summary
Researchers argue RoPE has intrinsic limits in long-context LLMs, especially for preserving order and token identity at great distances.
They proved and tested cases where attention can treat a moved token and a different token at the same position almost the same.
This can break simple retrieval tasks, including needle-in-a-haystack-style tests.
The findings suggest advertised context windows may overstate real usable performance.
Future long-context models may need new ways to encode position and sequence order.
