Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

TL;DR AI
2 min readKey summary
Researchers introduced Harness-1, a 20B retrieval subagent trained with reinforcement learning inside a stateful search harness.
The harness externalizes memory and bookkeeping—candidate pools, curated evidence, verification records, and compressed observations—so the model can focus on search, selection, verification, and stopping.
Harness-1 achieved strong results across eight benchmarks and showed especially good transfer to held-out retrieval tasks.
The work suggests search agents improve when environment-side state handles working memory, making reinforcement learning more effective for new retrieval settings.
