Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
TL;DR AI
2 min readKey summary
Researchers introduced Claw-Anything, a benchmark for always-on AI assistants that tests long-horizon reasoning across services and devices.
The benchmark simulates months of noisy, real-world user activity to measure how well agents handle persistent, proactive assistance.
GPT-5.5 achieved only 34.5% pass@1, showing today’s models still struggle in these rich digital environments.
The authors also released a data-generation pipeline that builds 2,000 training environments and improved the base model by 23.7%.
