From Model Scaling to System Scaling: Scaling the Harness in Agentic AI

TL;DR AI
2 min readKey summary
An arXiv paper argues agentic AI should be advanced by “scaling the harness,” not only by building larger models.
The focus shifts to the execution layer around models: auditable, modular components for memory, routing, orchestration, and governance.
The paper calls for new harness-level benchmarks to measure long-horizon agent performance more realistically.
It introduces CheetahClaws, a Python-native reference harness, and compares it with Claude Code and OpenClaw.
