Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic

TL;DR AI
2 min readKey summary
Researchers propose Vision-TL-Action, a neuro-symbolic model that generates robot trajectories from multi-view images and temporal-logic goals.
The system combines task tokens, visual tokens, and the robot’s initial state using bidirectional cross-attention and a flow-matching generator.
A training-only grounding loss helps map observations to structured task goals without requiring object geometry at inference.
It outperforms an oracle-state baseline on Panda and performs close to it on AntMaze, with ablations confirming the value of semantic grounding.
