Switch language한국어
Back to the list

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

TL;DR AI

Key summary

2 min read
  1. LongDS benchmarks AI agents on long-horizon data analysis using 68 real Kaggle notebook tasks across six domains and 2,225 turns.

  2. Across five leading models, the best result was 48.45% accuracy, showing limited reliability in extended analytical workflows.

  3. Performance fell sharply as conversations progressed, with most mistakes caused by failures to track and revise analytical state.

  4. The findings highlight a major gap in current agents’ ability to support dependable, realistic data science work.

Read the original