Survey Finds 5 Cutting-Edge AIs Disagreed on Claims 67% of the Time
TL;DR AI
2 min readKey summary
Lenz tested 1,000 user-submitted claims across five frontier AI models and found disagreement in 672 cases.
Only 328 claims received unanimous ratings, showing that the models often judged the same claim differently.
Some claims split all five models, underscoring the inconsistency of AI fact-checking on real-world inputs.
Lenz said human-labeled answers are being used to study where model judgments diverge from people.



