explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What LightOnOCR-3 does
  • Benchmarks, with the caveats
  • Speed
  • How to try it
  • How it compares with other options
  • What this means for what you build or pay
  • Limits and open questions
  • Summary
  • Related reading
← Back to blog

explainx / blog

LightOnOCR-3: Apache 2.0 OCR in 0.8B, 1B and 4B, Ranked on ParseBench

OCR, Document AI, Open Source AI, LightOn, Benchmarks

Part of Open-Weight Models

LightOnOCR-3 adds grounding boxes, image descriptions and chart tables to 0.8B, 1B and 4B Apache 2.0 OCR models. Benchmarks, speed on H100, and the caveats.

Oct 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
LightOnOCR-3: Apache 2.0 OCR in 0.8B, 1B and 4B, Ranked on ParseBench

Document parsing is the quiet bottleneck in most RAG and agent pipelines. A model that reads a PDF badly poisons everything after it. On October 8, 2026, LightOn released LightOnOCR-3, a family of three small open models built to read pages and say where each thing sits on them.

The headline says LightOnOCR-3 "tops ParseBench for open-weight AI." That is true in LightOn's comparison table. It is less clean on the public leaderboard. This post covers what the models do, the benchmark numbers with their caveats, the speed data, how to run them, and where they fit next to Mistral OCR and other open options. We read the model card, LightOn's Hugging Face blog post, the GitHub repository and the ParseBench leaderboard. We did not run the models ourselves.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?Three OCR models: 0.8B, 1B and 4B, all Apache 2.0
What is new vs LightOnOCR-2?Grounding with boxes, image descriptions, chart-to-table extraction
Is it the top open model on ParseBench?In LightOn's table, yes: 75.1 (4B) and 74.6 (0.8B) vs 74.3 for Infinity Parser Pro
Is it first on the public leaderboard?No. We saw KDL-Frontier-Parser-nano at 76.36 listed above it
Which size?LightOn recommends the 4B for most tasks
Commercial use?Yes. Apache 2.0
How to serve it?vLLM, or Transformers v5

What LightOnOCR-3 does

The model card calls it "a new family of highly performant lightweight OCR models." Compared with the previous generation, LightOn says the models bring "significant improvements in speed and transcription quality" and add visual understanding. Three things changed.

  1. Grounding. Call the model with the prompt grounding and every block returns with a label and a bounding box. Coordinates are normalized to 0 to 1000.
  2. Image descriptions. Photos and illustrations get a short description. That makes them searchable in a retrieval pipeline.
  3. Chart extraction. Charts become an HTML table of their data points. The numbers inside a bar chart become fields you can query.

LightOnOCR-3 grounded output turning a document page into labeled blocks with bounding boxes for open-weight OCR

Figure: a page and its grounded output (title, text, image and chart blocks). Fictional example for illustration from LightOn's Hugging Face blog post, used with credit to LightOn.

With an empty prompt, the models work as before and return the page text as markdown. LightOn says switching from LightOnOCR-2 needs no change, because this is the default behavior.

The three sizes

table · 3 cols
VariantArchitectureRole in LightOn's words
LightOnOCR-3-4BQwen3.5 vision-languageBest OCR model, recommended for most tasks
LightOnOCR-3-1BLightOnOCR-2 architectureDrop-in upgrade for existing deployments
LightOnOCR-3-0.8BQwen3.5 vision-languageFast and efficient model

The 0.8B and 4B adopt the Qwen3.5 architecture. LightOn says this "simplifies integration with existing tools and brings a significant speed-up."

The label set

Each block carries a type. The model card lists labels such as text, title, list, header, footer, page_number, footnote, caption, formula, code, table, image, chart, header_image, footer_image and aside_text. A plus suffix marks a continuation block. That gives you enough structure to drop page furniture before you chunk text for retrieval.

Grounding costs tokens. The card says it adds about 25 percent output tokens over plain transcription: 1,433 versus 1,158 tokens per page on average for the 4B.

Benchmarks, with the caveats

LightOn compares against open-weight OCR models that can run locally. The figures below come from its blog post. They are vendor-run numbers.

ParseBench

ParseBench is a benchmark from the LlamaIndex team. Its site describes it as "2,000 human-verified pages, 169K deterministic test rules, 5 capability dimensions."

table · 7 cols
ModelSize (B)TablesChartsSemantic formattingVisual groundingOverall (5 cats)
LightOnOCR-3-4B4.083.866.167.668.375.1
LightOnOCR-3-0.8B0.884.564.766.768.274.6
Infinity Parser Pro35.186.461.359.174.974.3
LightOnOCR-3-1B1.084.857.364.561.771.4
Chandra 24.089.265.161.451.270.1
MistralOCR4.1n/a73.940.166.471.268.2

The 4B and the 0.8B rank first and second, by a narrow margin. The gap to Infinity Parser Pro is 0.8 points for the 4B. Infinity Parser Pro has 35.1B listed parameters, about nine times the 4B. That size gap is the real story.

The weak spots show in the table. Chandra 2 beats it on tables (89.2). Infinity Parser Pro and Mistral beat it on visual grounding. The 1B drops to 71.4 because of charts and visual grounding. LightOn says this "highlights the benefits of starting from pre-trained VLMs for visual categories."

The leaderboard caveat

We opened the public ParseBench leaderboard on Hugging Face on October 9. It listed KDL-Frontier-Parser-nano at 76.36, Infinity-Parser2-Pro at 74.3 and Infinity-Parser2-Flash at 73.25 in its top three. The LightOnOCR-3 models were not in the top ten we saw. The leaderboard may not have an entry for them yet. So "tops ParseBench" means "tops the open models LightOn chose to compare." It does not mean first place on the public board. Treat the claim as strong, not final.

olmOCR-Bench

On olmOCR-Bench, the 4B scores 86.3 overall. That is 1.3 points behind Infinity Parser Pro (87.6) and 0.5 above Chandra 2 (85.8). The 0.8B scores 85.5, and the 1B scores 84.5. MistralOCR4.1 scores 81.9 in this table. The 4B leads on ArXiv pages (91.1). Infinity Parser Pro stays stronger on old scans and headers and footers. The headline that the models "outperform Mistral OCR 4.1, Chandra-OCR-2 and dots.mocr" on olmOCR-Bench holds for the overall column in the table.

French documents

On fr-bench-pdf2md, a French-document benchmark, the 4B leads at 74.1. It is strongest on handwritten pages (46.1) and forms (49.6). Handwriting at 46.1 is still low in absolute terms. Do not expect clean handwriting reads.

The normalization caveat

LightOn writes openly about a measurement problem. All these benchmarks use edit distance. A footnote can be written as HTML, LaTeX or Unicode, and the scorer treats them as different. The authors say a "simple deterministic rewrite can therefore change a model's score noticeably without changing what it actually extracted." They apply normalization functions to every model's raw output before scoring and publish those functions. The overall scores shift a lot for all top models. Read the benchmark gaps of one or two points as noise, and test on your own pages.

Speed

LightOn measured serving speed on the same 512 olmOCR-Bench pages, one H100 per model, vLLM 0.30.0, with the guidellm load tool.

LightOnOCR-3 serving speed table on one H100 for open-weight OCR models in grounding mode

Figure: speed on the same 512 olmOCR-Bench pages. Source: LightOn's Hugging Face blog post, used with credit to LightOn.

Key points from the post:

  • Resolution sets throughput. The 0.8B and 4B reach their best olmOCR-Bench scores at 400 DPI with a 5 MP cap, about 4.8k image tokens per page. Rendering at 1540 pixels on the long side cuts image tokens by almost two thirds. Peak throughput rises 44 percent for the 0.8B (4.78 pages per second) and 66 percent for the 4B (3.36 pages per second).
  • The 1B has the lowest single-page latency, at 2.7 seconds.
  • Versus Chandra-OCR-2, built on the same Qwen3.5-4B base, LightOnOCR-3-4B produces 14 percent fewer tokens per page. At each model's olmOCR-Bench settings it is 19 percent faster on a single page and serves 21 percent more pages per second.

The tradeoff: speed at 1540 pixels costs some accuracy relative to the 400 DPI setting, since that setting gave the best scores. LightOn does not give the accuracy hit in the text we read. Test both settings on your documents.

Output tokens per page for LightOnOCR-3 versus Chandra-OCR-2 and Infinity-Parser2-Pro in grounding mode

Figure: mean output tokens per page over the same 512 pages. Source: LightOn's Hugging Face blog post.

LightOn states the 0.8B and 4B generate 9 to 14 percent fewer output tokens than Chandra-OCR-2 and Infinity-Parser2-Pro, even though their output includes boxes, image descriptions and chart data. The company notes the comparison covers complete outputs, not formatting overhead alone.

How to try it

The model cards give two routes.

vLLM

bash
vllm serve lightonai/LightOnOCR-3-1B \
    --limit-mm-per-prompt '{"image": 1}' --mm-processor-cache-gb 0 --no-enable-prefix-caching

Then call the OpenAI-compatible endpoint with an image and, for grounding, a text part that says grounding. Use temperature 0.2 and top_p 0.9 as in the card's example. Swap the model ID for the 4B or 0.8B.

Transformers

The 1B loads with the LightOnOcr classes in Transformers v5. Use LightOnOcrForConditionalGeneration and LightOnOcrProcessor. On Apple silicon the card picks mps with float32.

Preprocessing tips from the card

  • Render PDFs at 200 DPI, with a target longest dimension of 1540 px.
  • Keep the aspect ratio.
  • Use only the empty prompt or grounding. The card says other instructions are out of distribution.

The LightOnOCR repository adds a minimal client, a command-line tool, a viewer and the scripts to reproduce the benchmarks. It also converts the raw output format into other formats.

How it compares with other options

table · 4 cols
OptionWeightsStrengthCaveat
LightOnOCR-3 (0.8B to 4B)Open, Apache 2.0Small, grounded output, chart tablesHandwriting is weak, public leaderboard entry not yet seen
Mistral OCR 4 and 4.1Closed APIStrong visual grounding on ParseBenchPer-page API cost, no local run
Interfaze-1-LiteOpen, Apache 2.0OCR plus speech plus structured extractionNeeds an 80 GB-class GPU for the full model per reports
MinerU 3.4OpenPipeline for RAG and agentsA multi-stage pipeline, not one model
Baidu Unlimited OCRPer its postLong-horizon one-shot parsingDifferent design goal
Infinity Parser ProOpen, 35.1BTop old-scan and header resultsRoughly nine times the 4B in size

For a RAG pipeline that already screenshots pages, see PixelRAG. Grounded output from LightOnOCR-3 helps when you want text chunks that still point back to a page region.

What this means for what you build or pay

  1. Self-hosting got cheaper. A 0.8B model that ranks near the top of a document benchmark fits a single modest GPU. An Apache 2.0 license removes the per-page fee and the data-sharing concern.
  2. One model replaces a pipeline. Layout detection, OCR, chart reading and image captioning were separate stages. The grounding mode gives you all of them in one pass. Fewer stages means fewer places for errors to compound.
  3. Pick on your documents. Tables, charts, handwriting and French forms each rank the models differently. A one-point overall lead is not a reason to switch.
  4. Budget for the token overhead. Grounding adds roughly a quarter more output tokens. Use plain mode where you do not need boxes.

A short evaluation plan: collect 50 real pages including your worst scans. Run the 0.8B and 4B in both modes. Score tables and numbers by hand on a sample. Compare against your current parser on the same pages.

Limits and open questions

  • Vendor-run benchmarks. The tables come from LightOn's post and its normalization functions. The authors published the scripts so others can reproduce them.
  • Leaderboard status. The public ParseBench board did not show LightOnOCR-3 in its top ten when we looked.
  • Handwriting. Scores of 33 to 46 on French handwritten pages are low.
  • Download counts. The 1B card showed 11 downloads on October 9, so there is little community feedback yet.
  • Hacker News. We searched Hacker News for LightOnOCR on October 9 and found no story, so there are no developer reports to quote.

Summary

LightOnOCR-3 gives open-weight OCR a grounded, chart-aware output format in three small sizes. The 4B and 0.8B post the best ParseBench scores in LightOn's own comparison of open models. The 0.8B does it at about one forty-fourth the size of Infinity Parser Pro. Check the public leaderboard and your own pages before you rely on the ranking.

Related reading

  • Mistral OCR 4 and 4.1: bounding boxes and the OCR API
  • Interfaze-1-Lite: open weights for OCR, speech and extraction
  • MinerU 3.4 for document parsing in RAG and agents
  • Baidu Unlimited OCR
  • PixelRAG: visual RAG over screenshots

Primary sources: the LightOnOCR-3-4B model card, the LightOn blog post, the LightOnOCR repository, the ParseBench site and the ParseBench leaderboard. The HuggingNews summary is here.

Scores, model sizes and leaderboard positions are accurate as of October 9, 2026. Leaderboards change quickly, so re-check before you decide.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 9, 2026

Tencent Youtu-Parsing-Omni: A 5B Open Model That Parses Documents, Audio and Video

Tencent has published the weights of Youtu-Parsing-Omni, a compact 5B model that reads document pages, charts, geometry figures, audio and video and answers in one structured JSON envelope. explainx.ai walks through the benchmark table, how to run it, and the license clause to check first.

Oct 6, 2026

Interfaze-1-Lite: An Apache 2.0 Model for OCR, Speech and Structured Extraction

Interfaze, a Y Combinator company, released interfaze-1-lite on October 5, 2026: an open-weight model aimed at deterministic backend jobs such as document OCR, speech-to-text, classification and extraction, returning confidence scores and bounding boxes. It runs on one 80 GB GPU. Here is what it is, the numbers, and how to test it against your current pipeline.

Sep 10, 2026

A 4B Open-Source VLM Reportedly Beats Qwen 122B on GeoGuessr-Style Benchmarks

A new 4B-parameter open-source vision-language model reportedly outperforms Alibaba's much larger 122B-parameter Qwen model on GeoGuessr-style benchmarks — guessing a photo's real-world location from visual clues alone. explainx.ai covers why this specific benchmark exists, what it actually measures, why a much smaller model beating a 30x larger one is plausible rather than implausible, and what it means for anyone choosing a vision-language model.