COMAP: Co-Evolving World Models and Agent Policies for LLM Agents

TL;DR AI
2 min readKey summary
Researchers introduced COMAP, a closed-loop framework where a textual world model and an LLM agent improve each other over time.
The world model predicts feedback for candidate actions, the agent uses that feedback to refine decisions, and the model is updated from the resulting trajectories.
In benchmarks, COMAP outperformed baselines across planning, web navigation, embodied tasks, and tool-use settings.
The approach shows a path for LLM agents to get better without external rewards by jointly adapting their internal world understanding and policy.
