Switch language한국어
Back to the list

HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced HoF-Bench, a benchmark built from 95 real CVEs discovered by AISLE across eight repositories.

  2. They tested ten analyzer backbones with a strict blinded judging setup to compare AI vulnerability discovery on real bugs.

  3. A minimal LLM-based analyzer rediscovered up to 65 CVEs, while no frontier model succeeded in the study.

  4. The results suggest that smaller LLM-based systems can be effective, but performance varies by code language and model consistency.

Read the original