Switch language한국어
Back to the list

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Claw-Anything, a benchmark for always-on AI assistants that tests long-horizon reasoning across services and devices.

  2. The benchmark simulates months of noisy, real-world user activity to measure how well agents handle persistent, proactive assistance.

  3. GPT-5.5 achieved only 34.5% pass@1, showing today’s models still struggle in these rich digital environments.

  4. The authors also released a data-generation pipeline that builds 2,000 training environments and improved the base model by 23.7%.

Read the original