Switch language한국어
Back to the list

LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LLMEval-Logic, a new Chinese benchmark for testing LLM logical reasoning.

  2. It combines expert-audited natural-language questions, formal annotations verified by Z3, and rubric-based grading.

  3. A specially hardened subset was built through adversarial workflows to make the test much tougher.

  4. Evaluation of 14 frontier models showed limited performance on the hardest tasks, exposing major reasoning gaps.

Read the original