IBM Releases Granite 4.0 3B Vision: A New Vision Language Model for Enterprise-Grade Document Data Extraction

TL;DR AI
2 min readKey summary
IBM released Granite 4.0 3B Vision, a compact vision-language adapter built on Granite 4.0 Micro for high-fidelity enterprise document parsing.
The model uses tiled image encoding and DeepStack-style visual token injection, and was trained on chart and table extraction data like ChartNet.
IBM says it performs strongly on document-understanding benchmarks and ranks near the top of the 2–4B parameter class for structured extraction.
The launch reflects a broader move toward smaller, task-specific multimodal models that extract business data more accurately and efficiently.



