HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

TL;DR AI
2 min readKey summary
Researchers introduced HoF-Bench, a benchmark built from 95 real CVEs discovered by AISLE across eight repositories.
They tested ten analyzer backbones with a strict blinded judging setup to compare AI vulnerability discovery on real bugs.
A minimal LLM-based analyzer rediscovered up to 65 CVEs, while no frontier model succeeded in the study.
The results suggest that smaller LLM-based systems can be effective, but performance varies by code language and model consistency.
