← File Toolkit

Local Tesseract · Private temporary jobs

Make a scanned PDF searchable with OCR

Add an invisible searchable text layer while preserving page appearance, or download plain text and structured OCR data.

1. Upload PDFs

Add at least one PDF to continue.

Native text stays native

Page evidence decides whether OCR is required. Reliable digital text is not degraded or duplicated.

Working-image corrections

Orientation, deskew and contrast operate on a bounded temporary raster. Original colour pages remain untouched by default.

Transparent quality

Page strategies, languages, confidence, transformations and low-confidence warnings remain available in structured data.

About PDF OCR

Private local OCR for scanned PDF documents

The OCR workflow classifies each page, preserves reliable native text, and sends only pages that need recognition through local Tesseract. Orientation, deskew, contrast, language selection, confidence, and page ranges remain under your control.

OCR outputs and controls

  • Create a searchable PDF while preserving original visuals
  • Download UTF-8 plain text
  • Export words, lines, regions, coordinates, confidence, and warnings as JSON
  • Select OCR languages, page ranges, orientation, deskew, and preprocessing

How to OCR a PDF

  1. Upload one or more scanned or mixed PDFs.
  2. Choose output formats, languages, pages, and correction options.
  3. Run OCR and download the validated searchable PDF, text, or JSON.

Frequently asked questions

Does OCR change the appearance of my PDF?

The default searchable-PDF output preserves original page visuals and adds an invisible text layer.

Can I OCR only selected pages?

Yes. Enter a page range such as 1-3,5 before starting processing.

Is an external OCR provider used?

No. OCR runs locally with installed Tesseract language data.