Switch language한국어
Back to the list

Baidu Qianfan Team Releases Qianfan-OCR: A 4B-Parameter Unified Document Intelligence Model

TL;DR AI

Key summary

2 min read
  1. Baidu's Qianfan team released Qianfan-OCR, a 4B-parameter end-to-end vision-language model for unified document tasks.

  2. The model uses a Qianfan-ViT encoder, a cross-modal adapter, and Qwen3-4B as the language backbone with a 32K context window.

  3. Layout-as-Thought is an optional thinking phase that produces structured layout representations before final output.

  4. Qianfan-OCR topped several end-to-end benchmarks including OmniDocBench v1.5 (93.12) and achieved 1.024 PPS with W8A8 quantization on an A100.

Read the original