BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

TL;DR AI
2 min readKey summary
Researchers introduced BenHalluEval, the first dedicated benchmark for hallucination detection in Bengali large language models.
The framework spans question answering, code-mixed QA, summarization, and reasoning, with 12,000 GPT-5.4-generated hallucinated candidates across 12 hallucination types.
It evaluates seven LLMs and adds BenHalluScore, a metric designed to balance hallucination detection with false-positive penalties.
The benchmark fills a major gap for Bengali, a low-resource language that had not been systematically assessed for hallucinations before.
