← 返回论文检索
CVPR 2026

PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks

Cheng Cui, Yubo Zhang, Ting Sun, Xueqing Wang, Hongen Liu, Manhui Lin, Yue Zhang, Tingquan Gao, Changda Zhou, Jiaxuan Liu, Zelun Zhang, Jing Zhang, Jun Zhang, Yi Liu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

The advent of "OCR 2.0" and large-scale vision-language models (VLMs) has set new benchmarks in text recogni- tion. However, these unified architectures often come with significant computational demands, challenges in precise text localization within complex layouts, and a propen- sity for textual hallucinations. Revisiting the prevailing notion that model scale is the sole path to high accu- racy, this paper introduces PP-OCRv5, a meticulously op- timized, lightweight OCR system with merely 5 million pa- rameters. We demonstrate that PP-OCRv5 achieves per- formance competitive with many billion-parameter VLMs on standard OCR benchmarks, while offering superior lo- calization precision and reduced hallucinations. The cor- nerstone of our success lies not in architectural expansion but in a data-centric investigation. We systematically dis- sect the role of training data by quantifying three critical dimensions: data difficulty, data accuracy, and data diver- sity. Our extensive experiments reveal that with a sufficient volume of high-quality, accurately labeled, and diverse data, the performance ceiling for traditional, efficient two- stage OCR pipelines is far higher than commonly assumed. This work provides compelling evidence for the viability of lightweight, specialized models in the large-model era and offers practical insights into data curation for OCR. The source code and models are publicly available at https: //github.com/PaddlePaddle/PaddleOCR.