Tencent has released the weights of Youtu-Parsing-Omni, a 5-billion-parameter model that takes a document page, a chart, a geometry figure, an audio clip or a video and answers with one structured JSON object. The model card on Hugging Face dates the release to October 2026, with a vLLM plugin and inference examples in the Youtu-Parsing GitHub repository. The headline claim is a 96.96 overall score on OmniDocBench, the top row among the models Tencent chose to compare.
That number deserves context, and so does the license. This post explains what the model does, what the benchmark tables actually show, how to run it, and the clause in its license that will rule it out for some teams.
TL;DR: the questions people will ask first
| Question | Short answer |
|---|---|
| What is it? | A 5B omni-modal parser: documents, images, charts, geometry, audio and video to JSON |
| Who made it? | Tencent's Youtu team, building on the earlier Youtu-Parsing and Youtu-LLM work |
| Is it open? | Weights, a vLLM plugin and examples are public; the license file is custom, not Apache or MIT |
| Best result | 96.96 overall on OmniDocBench, per the model card |
| Biggest catch | The license says it is not intended for use within the European Union |
| Technical report? | Marked "coming soon"; evaluation code is also still to come |
| What do I need to run it? | Python 3.10 or newer and a CUDA GPU |
What does Youtu-Parsing-Omni actually do?
Most document models do one job: turn a page image into Markdown. Youtu-Parsing-Omni treats parsing as a family of tasks, each chosen with a --task flag that maps to a prompt in the repository. The output is always a single JSON envelope that covers both perception (what is on the page or in the clip) and cognition (a caption, a narrative or a report).
The model card lists eight task keys across seven parsing families:
| Task key | Input | What comes back |
|---|---|---|
document | Document page | Layout elements with bounding boxes, text, LaTeX, tables, Markdown charts, Mermaid flowcharts, reading order |
natural_image | Photo | Entities and text with boxes, tags, captions, a global description |
graphics_chart | Chart | A Markdown table, notes, a caption |
graphics_flowchart | Flowchart | Mermaid code and a caption |
graphics_geometric | Geometry figure | Points, lines, arcs, shapes, relations and measurements |
audio | Audio clip | Vocal and non-vocal segments with timestamps, speakers, speech recognition, scene captions, acoustic events |
natural_video | Video | Temporal segments with visual elements, actions, interactions, camera motion and the audio track |
textrich_video | Slides or lecture video | Segments with OCR and speech recognition plus a Markdown structured report of the whole video |

The practical appeal is a single schema. If you build a retrieval pipeline over mixed company material, such as scanned contracts, slide decks, recorded meetings and product videos, you normally stitch together an OCR model, a chart extractor and a speech recognizer, then reconcile their formats. A unified envelope lets one service and one downstream parser handle all of it. For the retrieval side of that pipeline, see our guide to designing RAG context injection and the embedding model roundup.
How does the model work under the hood?
Tencent's description is short, and the technical report is not out yet, so the architecture details are limited to what the model card and the Hugging Face metadata show.
- An omni encoder. The card says image, audio and interleaved audio-visual video (frames plus the audio track) go through one encoder into one model.
- A 5.3B-parameter checkpoint. The Hugging Face metadata lists about 5.33 billion BF16 parameters, a file size of roughly 10.7 GB, and a model class named
youtu_vitawith custom code, so loading requires trusting remote code in Transformers. - A lineage of earlier Youtu work. The page links three earlier papers: Youtu-Parsing (high-parallelism decoding for parsing), Youtu-VL (unified vision-language supervision) and Youtu-LLM (a lightweight agentic base model). The earlier Youtu-Parsing was a 2.5B document model, and it appears in the new benchmark table at 93.74.
- Prompt-driven output. A chat template switches into a parsing mode when
enable_parsingis set, which is how the task prompts steer the output format.
Because the report is missing, treat training data, decoding speed and context limits as unknown. The model card does not publish throughput or per-page latency numbers.
Benchmarks: where does it actually lead?
Tencent's results come from its own model card, run against two benchmarks.
OmniDocBench
On OmniDocBench, which scores page parsing, Youtu-Parsing-Omni lands at 96.96 overall. The nearest rows in the card's table:
| Model | Size | Overall |
|---|---|---|
| Youtu-Parsing-Omni | 5.0B | 96.96 |
| TeleOCR | 1.2B | 96.91 |
| OvisOCR2 | 0.8B | 96.47 |
| PaddleOCR-VL-1.6 | 0.9B | 96.34 |
| MinerU2.5-Pro | 1.2B | 95.75 |
| Gemini 3 Pro | n/a | 92.91 |
| GPT-5.2 | n/a | 86.59 |
Read the gap carefully. The lead over TeleOCR is 0.05 points, and TeleOCR is a 1.2B model that actually wins on text edit distance (0.0267 against 0.0271 for Youtu-Parsing-Omni) and table structure score. Youtu-Parsing-Omni's strengths in the table are the structure-only table metric (98.20) and reading order (0.1104, lower is better). On formulas it trails PaddleOCR-VL-1.6 (96.80 against 97.53). In plain terms: it is at the top of a tightly packed group, not far ahead of it, and it is four to six times larger than the specialists next to it.
The comparison is also selective by the vendor's own note: older versions of the same family and models scoring below 86 are omitted. Third-party reproduction will tell us more.

OmniParsingBench
The newer OmniParsingBench spans natural images, graphics, audio and video. Here the model is second to Gemini-3-Pro, which averages 77.38 against Youtu-Parsing-Omni's 75.06. The split is interesting:
- It wins on chart parsing (95.02 against 92.79 for Gemini-3-Pro) and audio (78.08 against 76.74).
- It trails clearly on natural images (62.84 against 69.73) and geometry (74.33 against 85.43).
- It is close on natural video (74.18 against 74.82) and text-rich video (75.62 against 76.80).
Tencent adds that these numbers use a corrected ground truth and an LLM judge (Qwen3-235B-A22B-Instruct-2507), so they compare with each other but not with scores computed on the original release. Several rival models, including GPT-5.4 and Qwen3.5-397B, are not evaluated on audio or video at all, which makes the "second only to Gemini" framing narrower than it sounds.
Chemistry and music scores
The card also reports symbolic recognition results. On ChemOCR (image to SMILES) it scores 74.63 average similarity against 74.95 for DeepSeek-OCR 2, but 54.15 on the strict Tanimoto at 1.0 measure, the best listed. On PDMX, a full-page music score task, its ABC character error rate is 22.73 against 23.30 for the specialized LEGATO model, while LEGATO is slightly better on the other two error measures. These are niche tasks, but they show the model was trained beyond ordinary business documents.
How do I run it?
The repository gives two installation modes, and they are not interchangeable.
git clone https://github.com/TencentCloudADP/youtu-parsing.git
cd youtu-parsing/youtu_parsing_omni
# one mode per Python environment (Python 3.10+, CUDA GPUs for inference)
bash scripts/setup_env.sh # vLLM serving plus plugin
bash scripts/setup_env.sh --transformers # plain Transformers inference
The vLLM environment pins transformers==5.2.0 and the Transformers environment pins 5.10.2, so keep them in separate virtual environments. The README ships task prompts in prompts/youtu_parsing_omni.json and example scripts that take a --task argument.
A sensible first test is to run the same invoice or paper page through the document task, then a screenshot of a chart through graphics_chart, and diff the Markdown against your current OCR output. Compare on your own hardest pages, because published benchmark pages are cleaner than real scans.
The license catch: not for use in the EU
This is the part many summaries will miss. The LICENSE file in the Hugging Face repository opens with a notice that Youtu-Parsing "IS NOT INTENDED FOR USE WITHIN THE EUROPEAN UNION," and its first clause adds a territorial limitation that prevails in the event of any conflict. The Hugging Face page lists the license only as the custom "youtu-parsing".

The practical reading, which is ours and not legal advice: if you or your customers operate in the EU, you should not assume you can deploy these weights, and you should ask counsel before building on them. This is a recurring pattern with Chinese lab releases and with some US ones; the open-weight label covers a wide range of terms. We covered the same tension in Tencent's HY3 open-source agentic model, and the closed-source versus local open-source guide lists the license checks worth doing before adopting any model.
How does it compare with other document parsers?
Our recent coverage of this category gives useful anchors:
- Baidu Unlimited-OCR focuses on long, multi-page documents in one shot, a different problem from Tencent's multi-modality goal.
- MinerU 3.4 is a pipeline-style toolkit for PDFs and Office files aimed at RAG and agents, with a permissive open-source posture and a smaller footprint.
- Cohere Parse 5 is a hosted API priced per thousand pages, so no GPUs to manage but no weights either.
The trade is size and scope against specialization. If you only need page OCR at scale, a sub-2B specialist in the same benchmark table is cheaper to serve and statistically level on OmniDocBench. If you need charts, audio, video and captions from one endpoint, Youtu-Parsing-Omni is the only open option in this group that claims all of them, subject to the license.
What should builders do with it this week?
- Check the license first. If the EU clause blocks your use case, stop here and look at the alternatives above.
- Test on your own data. Build a 50-page set from your real scans, including tables with merged cells and poor photographs, and compare edit distance and table structure against your current parser.
- Try the audio and video tasks only if you need them. The OmniParsingBench numbers for natural images and geometry trail Gemini-3-Pro, so do not assume parity across every modality.
- Wait for the technical report before committing to production. Evaluation code is also marked as coming, so the benchmark claims cannot yet be reproduced independently.
- Budget for a 5B model. At roughly 10.7 GB of BF16 weights it needs a real GPU, unlike the 0.9B to 1.2B parsers it sits next to.
Limitations of what we know
We have not run the model, so every score here is Tencent's own report from the model card. The comparison tables omit weaker models by design, the technical report is unpublished, and there is no independent reproduction yet. The release is brand new, so community feedback barely exists yet. We will add an update section when the report or third-party tests appear.
Bottom line
Youtu-Parsing-Omni is a credible step toward one small open model that reads everything a business archive contains. Its page-parsing score is at the top of a crowded group rather than ahead of it, its audio and video results are strong but unproven outside Tencent's own benchmark, and its license excludes the EU. Treat it as a strong candidate to test, not a default to adopt.
Related reading
- Baidu Unlimited-OCR: one-shot long-horizon document parsing
- MinerU 3.4: PDF and Office parsing for RAG and agents
- Cohere Parse 5: near-frontier document parsing
- Tencent HY3: 295B open-source agentic model
- RAG context injection pipeline design
- Top 10 open and closed embedding models
- Closed-source AI versus local open-source alternatives
Model details, scores and license terms are accurate as of October 9, 2026 and come from the Hugging Face model card and the GitHub README. Check both for changes.
