MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
TL;DR AI
2 min readKey summary
MiroEval presents a benchmark and evaluation framework for deep research systems that measures both process and outcome.
The benchmark includes 100 real-user tasks, with 70 text-only and 30 multimodal items, and uses a dual-path pipeline for updates.
Evaluation across 13 systems shows process quality predicts outcomes and multimodal tasks reduce performance by 3–10 points.
