AI models often give the right answers but point to the wrong sources

TL;DR AI
2 min readKey summary
CiteVQA is a new benchmark that evaluates both answer correctness and whether a model cites the exact supporting location in a document.
Testing 20 multimodal models found a wide gap: many systems could answer questions correctly but still failed to point to the right page, paragraph, table, or figure.
The study suggests source-location retrieval is a major bottleneck, and stronger retrieval significantly improves performance.
For law, finance, and medicine, the finding matters because an answer is only reliable if it can be traced to the right evidence.
