Use the following Docker images depending on your GPU platform:
Default (CUDA 13.0):
docker pull vllm/vllm-openai:unlimited-ocr
For Hopper GPUs (CUDA 12.9)
docker pull vllm/vllm-openai:unlimited-ocr-cu129
SGLang
Set up the environment (uv-managed virtualenv). Install the local SGLang wheel first,
then pin kernels==0.9.0 and install PyMuPDF for PDF-to-image conversion:
--model_dir baidu/Unlimited-OCR # Local path or Hugging Face model ID
--gpu 0 # CUDA_VISIBLE_DEVICES value
--server_log ./log/sglang_server.log
For OmniDocBench evaluation, you need to perform the following post-processing.
DET_RE = re.compile(r'<\|det\|>([^<\s]+)(?:\s*\[[^\]]*\])?\s*<\|/det\|>(.*)', re.DOTALL)
def remove_det(raw: str) -> str:
"""
Strip <|det|>type [bbox]<|/det|> markers, group lines belonging to the
same block with \n, and separate different blocks with \n\n.
"""
blocks = []
cur = None
for line in raw.splitlines():
line = line.rstrip()
if not line:
continue
m = DET_RE.match(line)
if m:
category, content = m.group(1).strip(), m.group(2).strip()
if category == 'image':
continue
if cur is not None:
blocks.append(cur)
cur = [content] if content else []
continue
if cur is None:
cur = []
cur.append(line)
if cur is not None:
blocks.append(cur)
text = '\n\n'.join('\n'.join(b) for b in blocks).strip()
return text
@misc{yin2026unlimitedocrworks,
title={Unlimited OCR Works},
author={Youyang Yin and Huanhuan Liu and YY and Qunyi Xie and Chaorun Liu and Shiqi Yang and Shaohua Wang and Zhanlong Liu and Hao Zou and Jinyue Chen and Shu Wei and Jingjing Wu and Mingxin Huang and Zhen Wu and Guibin Wang and Tengyu Du and Lei Jia},
year={2026},
eprint={2606.23050},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.23050},
}
Unlimited OCR Works
Welcome the Era of One-shot Long-horizon Parsing.
Release
Inference
Transformers
Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.3 + CUDA12.9:
vLLM
Please refer to the official vLLM recipe for deployment details:
Recipe: https://recipes.vllm.ai/baidu/Unlimited-OCR
Docker Images
Use the following Docker images depending on your GPU platform:
Default (CUDA 13.0):
For Hopper GPUs (CUDA 12.9)
SGLang
Set up the environment (uv-managed virtualenv). Install the local SGLang wheel first, then pin
kernels==0.9.0and install PyMuPDF for PDF-to-image conversion:Start the SGLang server:
Send streaming requests to the OpenAI-compatible API:
For batch inference,
infer.pystarts the SGLang server automatically and sends concurrent requests for an image directory or PDF:Useful options:
For OmniDocBench evaluation, you need to perform the following post-processing.
Visualization
Acknowledgement
We would like to thank Deepseek-OCR, Deepseek-OCR-2, PaddleOCR for their valuable models and ideas.
Citation