Can AI agents conduct open-ended AI research? Early evidence from two case studies
TL;DR AI
2 min readKey summary
A shadow evaluation study tested frontier AI agents on two unpublished NeurIPS 2026 papers, with the original authors grading the results.
Over six days of compute, the agents completed engineering tasks without human help but did not make meaningful progress on the core research questions.
Both papers were ultimately rejected, and the study highlighted recurring failure modes such as poor judgment, weak backtracking, limited resource awareness, instruction drift, and uncreative responses to design flaws.
The findings provide early empirical evidence on a key question for AI R&D automation: whether agents can handle open-ended AI research.
