explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people will ask first
  • What does Youtu-Parsing-Omni actually do?
  • How does the model work under the hood?
  • Benchmarks: where does it actually lead?
  • How do I run it?
  • The license catch: not for use in the EU
  • How does it compare with other document parsers?
  • What should builders do with it this week?
  • Limitations of what we know
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Tencent Youtu-Parsing-Omni: A 5B Open Model That Parses Documents, Audio and Video

Document Parsing, Tencent, Open Source AI, Multimodal AI, OCR

Part of Open-Weight Models

Tencent released Youtu-Parsing-Omni, a 5B open-weight model that turns documents, charts, audio and video into JSON. Scores, license catch and setup.

Oct 9, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Tencent Youtu-Parsing-Omni: A 5B Open Model That Parses Documents, Audio and Video

Tencent has released the weights of Youtu-Parsing-Omni, a 5-billion-parameter model that takes a document page, a chart, a geometry figure, an audio clip or a video and answers with one structured JSON object. The model card on Hugging Face dates the release to October 2026, with a vLLM plugin and inference examples in the Youtu-Parsing GitHub repository. The headline claim is a 96.96 overall score on OmniDocBench, the top row among the models Tencent chose to compare.

That number deserves context, and so does the license. This post explains what the model does, what the benchmark tables actually show, how to run it, and the clause in its license that will rule it out for some teams.

TL;DR: the questions people will ask first

table · 2 cols
QuestionShort answer
What is it?A 5B omni-modal parser: documents, images, charts, geometry, audio and video to JSON
Who made it?Tencent's Youtu team, building on the earlier Youtu-Parsing and Youtu-LLM work
Is it open?Weights, a vLLM plugin and examples are public; the license file is custom, not Apache or MIT
Best result96.96 overall on OmniDocBench, per the model card
Biggest catchThe license says it is not intended for use within the European Union
Technical report?Marked "coming soon"; evaluation code is also still to come
What do I need to run it?Python 3.10 or newer and a CUDA GPU
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What does Youtu-Parsing-Omni actually do?

Most document models do one job: turn a page image into Markdown. Youtu-Parsing-Omni treats parsing as a family of tasks, each chosen with a --task flag that maps to a prompt in the repository. The output is always a single JSON envelope that covers both perception (what is on the page or in the clip) and cognition (a caption, a narrative or a report).

The model card lists eight task keys across seven parsing families:

table · 3 cols
Task keyInputWhat comes back
documentDocument pageLayout elements with bounding boxes, text, LaTeX, tables, Markdown charts, Mermaid flowcharts, reading order
natural_imagePhotoEntities and text with boxes, tags, captions, a global description
graphics_chartChartA Markdown table, notes, a caption
graphics_flowchartFlowchartMermaid code and a caption
graphics_geometricGeometry figurePoints, lines, arcs, shapes, relations and measurements
audioAudio clipVocal and non-vocal segments with timestamps, speakers, speech recognition, scene captions, acoustic events
natural_videoVideoTemporal segments with visual elements, actions, interactions, camera motion and the audio track
textrich_videoSlides or lecture videoSegments with OCR and speech recognition plus a Markdown structured report of the whole video

A document sheet unfolding into tidy rows, illustrating how a document parsing model turns pages into structured output

The practical appeal is a single schema. If you build a retrieval pipeline over mixed company material, such as scanned contracts, slide decks, recorded meetings and product videos, you normally stitch together an OCR model, a chart extractor and a speech recognizer, then reconcile their formats. A unified envelope lets one service and one downstream parser handle all of it. For the retrieval side of that pipeline, see our guide to designing RAG context injection and the embedding model roundup.

How does the model work under the hood?

Tencent's description is short, and the technical report is not out yet, so the architecture details are limited to what the model card and the Hugging Face metadata show.

  • An omni encoder. The card says image, audio and interleaved audio-visual video (frames plus the audio track) go through one encoder into one model.
  • A 5.3B-parameter checkpoint. The Hugging Face metadata lists about 5.33 billion BF16 parameters, a file size of roughly 10.7 GB, and a model class named youtu_vita with custom code, so loading requires trusting remote code in Transformers.
  • A lineage of earlier Youtu work. The page links three earlier papers: Youtu-Parsing (high-parallelism decoding for parsing), Youtu-VL (unified vision-language supervision) and Youtu-LLM (a lightweight agentic base model). The earlier Youtu-Parsing was a 2.5B document model, and it appears in the new benchmark table at 93.74.
  • Prompt-driven output. A chat template switches into a parsing mode when enable_parsing is set, which is how the task prompts steer the output format.

Because the report is missing, treat training data, decoding speed and context limits as unknown. The model card does not publish throughput or per-page latency numbers.

Benchmarks: where does it actually lead?

Tencent's results come from its own model card, run against two benchmarks.

OmniDocBench

On OmniDocBench, which scores page parsing, Youtu-Parsing-Omni lands at 96.96 overall. The nearest rows in the card's table:

table · 3 cols
ModelSizeOverall
Youtu-Parsing-Omni5.0B96.96
TeleOCR1.2B96.91
OvisOCR20.8B96.47
PaddleOCR-VL-1.60.9B96.34
MinerU2.5-Pro1.2B95.75
Gemini 3 Pron/a92.91
GPT-5.2n/a86.59

Read the gap carefully. The lead over TeleOCR is 0.05 points, and TeleOCR is a 1.2B model that actually wins on text edit distance (0.0267 against 0.0271 for Youtu-Parsing-Omni) and table structure score. Youtu-Parsing-Omni's strengths in the table are the structure-only table metric (98.20) and reading order (0.1104, lower is better). On formulas it trails PaddleOCR-VL-1.6 (96.80 against 97.53). In plain terms: it is at the top of a tightly packed group, not far ahead of it, and it is four to six times larger than the specialists next to it.

The comparison is also selective by the vendor's own note: older versions of the same family and models scoring below 86 are omitted. Third-party reproduction will tell us more.

Two circular check impressions beside a stamp, representing reproducible benchmark results for the OmniDocBench comparison

OmniParsingBench

The newer OmniParsingBench spans natural images, graphics, audio and video. Here the model is second to Gemini-3-Pro, which averages 77.38 against Youtu-Parsing-Omni's 75.06. The split is interesting:

  • It wins on chart parsing (95.02 against 92.79 for Gemini-3-Pro) and audio (78.08 against 76.74).
  • It trails clearly on natural images (62.84 against 69.73) and geometry (74.33 against 85.43).
  • It is close on natural video (74.18 against 74.82) and text-rich video (75.62 against 76.80).

Tencent adds that these numbers use a corrected ground truth and an LLM judge (Qwen3-235B-A22B-Instruct-2507), so they compare with each other but not with scores computed on the original release. Several rival models, including GPT-5.4 and Qwen3.5-397B, are not evaluated on audio or video at all, which makes the "second only to Gemini" framing narrower than it sounds.

Chemistry and music scores

The card also reports symbolic recognition results. On ChemOCR (image to SMILES) it scores 74.63 average similarity against 74.95 for DeepSeek-OCR 2, but 54.15 on the strict Tanimoto at 1.0 measure, the best listed. On PDMX, a full-page music score task, its ABC character error rate is 22.73 against 23.30 for the specialized LEGATO model, while LEGATO is slightly better on the other two error measures. These are niche tasks, but they show the model was trained beyond ordinary business documents.

How do I run it?

The repository gives two installation modes, and they are not interchangeable.

bash
git clone https://github.com/TencentCloudADP/youtu-parsing.git
cd youtu-parsing/youtu_parsing_omni

# one mode per Python environment (Python 3.10+, CUDA GPUs for inference)
bash scripts/setup_env.sh                 # vLLM serving plus plugin
bash scripts/setup_env.sh --transformers  # plain Transformers inference

The vLLM environment pins transformers==5.2.0 and the Transformers environment pins 5.10.2, so keep them in separate virtual environments. The README ships task prompts in prompts/youtu_parsing_omni.json and example scripts that take a --task argument.

A sensible first test is to run the same invoice or paper page through the document task, then a screenshot of a chart through graphics_chart, and diff the Markdown against your current OCR output. Compare on your own hardest pages, because published benchmark pages are cleaner than real scans.

The license catch: not for use in the EU

This is the part many summaries will miss. The LICENSE file in the Hugging Face repository opens with a notice that Youtu-Parsing "IS NOT INTENDED FOR USE WITHIN THE EUROPEAN UNION," and its first clause adds a territorial limitation that prevails in the event of any conflict. The Hugging Face page lists the license only as the custom "youtu-parsing".

A badge with an open padlock and green ring illustrating the license terms attached to open-weight models

The practical reading, which is ours and not legal advice: if you or your customers operate in the EU, you should not assume you can deploy these weights, and you should ask counsel before building on them. This is a recurring pattern with Chinese lab releases and with some US ones; the open-weight label covers a wide range of terms. We covered the same tension in Tencent's HY3 open-source agentic model, and the closed-source versus local open-source guide lists the license checks worth doing before adopting any model.

How does it compare with other document parsers?

Our recent coverage of this category gives useful anchors:

  • Baidu Unlimited-OCR focuses on long, multi-page documents in one shot, a different problem from Tencent's multi-modality goal.
  • MinerU 3.4 is a pipeline-style toolkit for PDFs and Office files aimed at RAG and agents, with a permissive open-source posture and a smaller footprint.
  • Cohere Parse 5 is a hosted API priced per thousand pages, so no GPUs to manage but no weights either.

The trade is size and scope against specialization. If you only need page OCR at scale, a sub-2B specialist in the same benchmark table is cheaper to serve and statistically level on OmniDocBench. If you need charts, audio, video and captions from one endpoint, Youtu-Parsing-Omni is the only open option in this group that claims all of them, subject to the license.

What should builders do with it this week?

  1. Check the license first. If the EU clause blocks your use case, stop here and look at the alternatives above.
  2. Test on your own data. Build a 50-page set from your real scans, including tables with merged cells and poor photographs, and compare edit distance and table structure against your current parser.
  3. Try the audio and video tasks only if you need them. The OmniParsingBench numbers for natural images and geometry trail Gemini-3-Pro, so do not assume parity across every modality.
  4. Wait for the technical report before committing to production. Evaluation code is also marked as coming, so the benchmark claims cannot yet be reproduced independently.
  5. Budget for a 5B model. At roughly 10.7 GB of BF16 weights it needs a real GPU, unlike the 0.9B to 1.2B parsers it sits next to.

Limitations of what we know

We have not run the model, so every score here is Tencent's own report from the model card. The comparison tables omit weaker models by design, the technical report is unpublished, and there is no independent reproduction yet. The release is brand new, so community feedback barely exists yet. We will add an update section when the report or third-party tests appear.

Bottom line

Youtu-Parsing-Omni is a credible step toward one small open model that reads everything a business archive contains. Its page-parsing score is at the top of a crowded group rather than ahead of it, its audio and video results are strong but unproven outside Tencent's own benchmark, and its license excludes the EU. Treat it as a strong candidate to test, not a default to adopt.

Related reading

  • Baidu Unlimited-OCR: one-shot long-horizon document parsing
  • MinerU 3.4: PDF and Office parsing for RAG and agents
  • Cohere Parse 5: near-frontier document parsing
  • Tencent HY3: 295B open-source agentic model
  • RAG context injection pipeline design
  • Top 10 open and closed embedding models
  • Closed-source AI versus local open-source alternatives

Model details, scores and license terms are accurate as of October 9, 2026 and come from the Hugging Face model card and the GitHub README. Check both for changes.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 9, 2026

LightOnOCR-3: Apache 2.0 OCR in 0.8B, 1B and 4B, Ranked on ParseBench

LightOn released LightOnOCR-3 on October 8, 2026. Three Apache 2.0 sizes read pages, return labeled bounding boxes, describe images and turn charts into tables. The company says the 4B and 0.8B lead open models on ParseBench. The public leaderboard tells a more mixed story.

Aug 11, 2026

Qwen-MM-Plugins: Make Claude Code, Codex and OpenClaw Multimodal

Qwen shipped a plugin suite on August 10, 2026 that bolts multimodal capability onto agent harnesses it doesn't own — Claude Code, Codex, Gemini CLI, OpenClaw and more. Eight capabilities, Apache-2.0, skills plus on-demand MCP servers. The catch is a DashScope API key.

Jun 29, 2026

Gemma 4 31B on Cerebras: 1,800+ TPS — The Fastest Multimodal Inference Yet

Google DeepMind's Gemma 4 31B hits 1,851 TPS on Cerebras — first multimodal model at wafer-scale speed. Haiku 4.5-class intelligence, 18× faster, public preview now.