Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents

TL;DR AI
2 min readKey summary
Co-ReAct introduces rubric-guided step-level planning for ReAct agents, injecting rubrics at each inference step to steer reasoning and actions.
A rubric generator is trained with GRPO using list-wise Spearman correlation against expert rankings, helping it learn better next-step guidance.
The approach improves search-intensive reasoning and action selection on benchmarks such as DeepResearchBench and SQA-CS-V2.
Results show gains over ReAct and other baselines across both open and closed models, without changing the agent’s core decision logic.
