Switch language한국어
Back to the list

New math benchmark reveals AI models confidently solve problems that have no solution

TL;DR AI

Key summary

2 min read
  1. SOOHAK, a new 439-problem math benchmark, tests both graduate-level problem solving and intentionally flawed refusal cases.

  2. Top frontier models performed well below human experts on the hardest tasks, showing limits in advanced mathematical reasoning.

  3. No model scored above 50% on spotting unsolvable or invalid questions, exposing weak refusal and error-detection behavior.

  4. The results suggest that scaling AI improves more routine math, but not yet reliable judgment about when a problem has no valid answer.

Read the original