Built by Thiago Lima and open-sourced as ocr-it on GitHub, it targets documents where text is visible but not selectable — Kindle readers, slide decks, locked PDF viewers. Tesseract runs in an offscreen document with no outbound network requests and no install-time host permissions. Auto-pagination advances pages until duplicate text or a cap stops the run.