Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

TL;DR AI
2 min readKey summary
A study of 16 language models across eight families and 800 tasks found that hidden-state similarity often does not match shared reasoning.
Models looked more alike on problems they answered incorrectly, suggesting representational similarity can rise when reasoning fails.
Pre-decision states were similar, but post-decision states diverged, showing that internal trajectories split after choosing an answer.
Although some shared information was decodable, it had little causal impact on outputs, challenging simple interpretability claims and ensemble assumptions.
