INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

TL;DR AI
2 min readKey summary
Researchers released INS-ActBench, a benchmark for evaluating actuarial skills in large language models.
Built from public exams and sample questions from 16 actuarial associations, it contains 12,050 questions across knowledge, long-context case reasoning, and spreadsheet/R tasks.
Across nine LLMs and actuarial experts, models performed well on standardized knowledge but much worse on case reasoning and tool-based workflows.
The results highlight a gap between exam-style success and real-world actuarial reliability, especially in context-heavy and jurisdiction-specific work.
