Switch language한국어
Back to the list

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LaRA, a layer-wise method for detecting data contamination in RL-post-trained language models.

  2. LaRA measures perturbation sensitivity, directional collapse, and local rigidity to spot contamination through representation changes, not output signals.

  3. The paper finds contamination creates systematic layer-by-layer deviations and LaRA outperforms prior baselines on RL-trained reasoning models.

  4. The approach could make leakage detection more reliable and strengthen trust in evaluation and generalization claims.

Read the original