CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning
TL;DR AI
2 min readKey summary
Researchers introduced CopT, a new LLM reasoning framework that drafts an answer first and then reflects only when needed.
CopT uses contrastive checks between discrete and continuous inputs to estimate answer reliability and decide whether to keep thinking.
The method reportedly improves accuracy while reducing token usage across reasoning benchmarks, including math, coding, and agentic tasks.
Because it is training-free, CopT could lower inference cost and latency while making model outputs more reliable on hard problems.
