Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

TL;DR AI
2 min readKey summary
QuestBench is a new course-based benchmark where students design and review questions to test AI systems.
The project produced 256 questions across 14 humanities and social-science areas, creating a practical AI evaluation dataset.
Across 13 systems, performance was weak: the best system passed 57.58% of questions, while the average was 16.85%.
The work shifts AI education from productivity use toward judging the reliability of machine-generated knowledge.
