LLMs as Annotators of Credibility Assessment in Danish Asylum Decisions: Evaluating Classification Performance and Errors Beyond Aggregated Metrics

TL;DR AI
2 min readKey summary
Researchers introduced RAB-Cred, a new dataset for annotating credibility assessments in Danish asylum decisions.
They benchmarked 21 open-weight LLMs across 30 prompt setups to detect whether credibility assessments appear and whether they are positive or negative.
Results show the models can be useful, but performance varies widely by model and prompt, with notable confusion patterns and inconsistent errors.
The study ties mistakes to annotator confidence and case difficulty, suggesting LLMs could lower annotation costs in low-resource legal NLP, but need careful validation before deployment.
