LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
TL;DR AI
2 min readKey summary
Researchers introduced LLMEval-Logic, a new Chinese benchmark for testing LLM logical reasoning.
It combines expert-audited natural-language questions, formal annotations verified by Z3, and rubric-based grading.
A specially hardened subset was built through adversarial workflows to make the test much tougher.
Evaluation of 14 frontier models showed limited performance on the hardest tasks, exposing major reasoning gaps.
