Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

TL;DR AI
2 min readKey summary
Researchers proposed ESPP, a persona-panel method for evaluating generative UI outputs more like a diverse group of users than a single LLM judge.
Each simulated persona rates a UI screenshot, exchanges opinions through a bounded-confidence process, and is then aggregated into a final judgment.
In experiments, ESPP aligned much better with human evaluations and outperformed a prompt-ensemble baseline.
The approach also surfaces subgroup disagreements that a homogeneous evaluator would likely miss.
