Switch language한국어
Back to the list

Do Audio-Visual Large Language Models Really See and Hear?

TL;DR AI

Key summary

2 min read
  1. Researchers performed the first mechanistic interpretability study of Audio-Visual Large Language Models (AVLLMs).

  2. They found rich audio semantics encoded in intermediate layers, but these signals often fail to influence final text when audio conflicts with vision.

  3. Probing showed deeper fusion layers disproportionately favor visual representations, and the imbalance traces to limited audio-specific alignment during training.

Read the original