SketchVLM: Vision language models can annotate images to explain thoughts and guide users
TL;DR AI
2 min readKey summary
SketchVLM is a training-free, model-agnostic framework that lets vision-language models explain their answers with editable SVG sketches on images.
The system prompts VLMs to produce non-destructive overlays, making reasoning easier to inspect and refine directly on the image.
Across seven benchmarks, SketchVLM improved both task accuracy and annotation quality over baselines.
It also supports single-turn and multi-turn human-AI interaction, working with models such as Gemini-3-Pro and GPT-5.
