Zhipu AI's GLM-5V-Turbo turns design mockups directly into executable front-end code

Key summary
GLM-5V-Turbo is Zhipu AI's first multimodal coding base model and handles images, video, and text.
It is designed for agent workflows and can convert design mockups into executable front-end code.
The model integrates with agents like Claude Code and OpenClaw and offers thinking mode, streaming output, function calling, and context caching.
Technical specs include a 200,000-token context window, up to 128,000-token output, parallel multi-token prediction, and a new CogViT vision encoder.
Training and tooling combine RL across 30+ task types, agentic meta-skills in pre-training, a multi-level verifiable data system, a multimodal toolchain with box-drawing/screenshot/website-reading tools, and claimed leading results and strong AndroidWorld and WebVoyager benchmark scores.



