Baidu Open Sources End-to-End Long Document OCR Model Unlimited-OCR
Baidu open sources a new OCR model, Unlimited OCR, which specializes in parsing dozens of pages of long documents at once. It achieves a new SOTA on OmniDocBench with a comprehensive score of 93.23%, surpassing DeepSeek OCR. The model's core innovation is the Reference Sliding Window Attention (R-SWA) mechanism. Through a "soft forgetting" strategy, it keeps the KV Cache constant, ensuring inference speed does not increase with document length. At 6000 Tokens, TPS improves by approximately 35%.
Baidu open sources a new OCR model, Unlimited OCR, which specializes in parsing dozens of pages of long documents at once. It achieves a new SOTA on OmniDocBench with a comprehensive score of 93.23%, surpassing DeepSeek OCR. The model's core innovation is the Reference Sliding Window Attention (R-SWA) mechanism. Through a "soft forgetting" strategy, it keeps the KV Cache constant, ensuring inference speed does not increase with document length. At 6000 Tokens, TPS improves by approximately 35%.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.