A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
TL;DR AI
2 min readKey summary
Researchers introduced TASTE, a method that automatically generates agent benchmark tasks by evolving valid tool sequences and refining them into harder problems.
Using TASTE, they created τ^c-Bench, a new benchmark with broader tool-use coverage and more diverse task combinations.
Agents that were near saturation on τ^2-Bench performed much worse on τ^c-Bench, suggesting current benchmarks can overstate progress.
The work highlights a scalable path to build tougher, more realistic evaluations for future tool-using models.
