Scan → Searchable PDF

Make Scanned PDFs Searchable (and Bates-Numbered)

For filings, discovery and archives: OhMyOCR renders each scanned page, runs OCR, lets you verify and correct low-confidence text, then rebuilds a two-layer PDF — original image on top, invisible corrected text underneath — with optional sequential Bates stamps.

No credit card Free credits Visual verification

How it works

  1. 1

    Upload scanned PDFs or images — batches queue automatically and process in order.

  2. 2

    OCR extracts the text with per-word coordinates and confidence; risky regions are flagged for review so misreads never hide in the text layer.

  3. 3

    Export a merged two-layer searchable PDF, with Bates numbering stamped continuously across the batch if you need it.

Why OhMyOCR

The text layer is corrected, not raw

Unlike one-click OCR tools, the searchable layer includes your corrections — the PDF you produce searches on what the document actually says.

Bates numbering across a batch

Sequential stamps (with your prefix) are applied across every page of every file in the export — continuous numbering for productions and filings.

Low-confidence text is visible

A search that silently misses one “not” can change a case. Low-confidence regions are flagged with image snippets so a human clears them before the PDF is built.

Zero-LLM option and auto-deletion

A deterministic no-AI extraction mode and per-account auto-deletion of originals N days after processing support confidentiality requirements.

A searchable PDF is only as good as its hidden text

Every OCR tool can stamp a text layer under a scan. The problem is what is in that layer: raw OCR output, errors included. A production set where “not” was misread in one hot document, or a name OCR’d into a variant spelling, fails silently — the search runs, returns nothing, and nobody knows the document was ever missed. In discovery and records work, that invisible failure is the expensive one.

OhMyOCR builds the text layer from reviewed text. Low-confidence regions are flagged with image snippets before export; a human clears them (keyboard-only, J/K/Enter), and the corrected reading — not the raw guess — is what goes under the image. The result looks identical to the scan and searches on what the document actually says.

Bates numbering and batch production

The Export Center merges an entire batch into one two-layer PDF with sequential Bates stamps — your prefix, continuous numbering across every page of every file, applied as a legible corner stamp. Combine it with the verified-review pass and you have a defensible pipeline: scan in, flags cleared, corrected text layer, numbered pages out.

For confidentiality-sensitive work: extraction offers a deterministic zero-LLM mode, originals can auto-delete N days after processing, and documents are never used for model training. (We do not currently offer HIPAA BAAs — medical records custodians should plan accordingly.)

Frequently asked questions

What is a two-layer (sandwich) PDF?

A PDF where each page shows the original scanned image with an invisible text layer underneath. It looks identical to the scan but supports search, select and copy — the standard for e-filing and archives.

Can I add Bates numbers?

Yes. Enable Bates numbering in the Export Center, set your prefix, and stamps run continuously across all files in the batch, applied to the page corner.

Why review before exporting?

Because a searchable PDF is only as good as its hidden text: if OCR misread a name, searches for that name silently fail. Flagged regions let you fix exactly those words before the text layer is built.

Does it handle photographs of documents, not just scans?

Yes — phone photos work. Pages are rendered and OCR runs with coordinates either way.

Is this suitable for confidential documents?

Documents are processed through our own pipeline, never used for training, with optional auto-deletion of originals. Note: we do not currently offer HIPAA BAAs.

Try it on your own file

Free to start, no credit card, results you can verify line by line.

Get started — free