Baidu Qianfan Team Releases Qianfan-OCR: A 4B-Parameter Unified Document Intelligence Model

TL;DR AI
2 min readKey summary
Baidu's Qianfan team released Qianfan-OCR, a 4B-parameter end-to-end vision-language model for unified document tasks.
The model uses a Qianfan-ViT encoder, a cross-modal adapter, and Qwen3-4B as the language backbone with a 32K context window.
Layout-as-Thought is an optional thinking phase that produces structured layout representations before final output.
Qianfan-OCR topped several end-to-end benchmarks including OmniDocBench v1.5 (93.12) and achieved 1.024 PPS with W8A8 quantization on an A100.



