Skip to content

PDF OCR: Extract Text from Scanned PDFs

xFormat.space runs OCR on scanned PDFs inside your browser, rendering each page and reading it with an on-device recognition model, so you get page-by-page text without uploading the document.

Runs in your browser — files never leave your device.

How it works

If you can't select or copy text in a PDF, it's probably a scan, which is just pictures of pages. This tool renders each page, recognizes the text on it locally on your device and returns plain text you can copy or save. The PDF never leaves your computer.

  1. 1

    Drop a scanned PDF onto the page and open OCR.

  2. 2

    Choose the language of the document.

  3. 3

    Click Start Recognition, wait for every page to finish, then copy the text or download it as a TXT file.

What you get, and what you don't

Output is plain text, with each page marked by a "--- Page N ---" line, plus a confidence score and a count of text regions. It is not a searchable PDF: the tool does not write a hidden text layer back into the file, so your original PDF stays as it was. If your PDF already lets you select text, you don't need OCR at all; just copy from it.

Scanned PDF accuracy tips

Pages are rendered at 2x scale before recognition, so the quality of the scan sets the ceiling. A scan around 300 dpi, upright and high-contrast, reads far better than a skewed, grey or low-resolution one. Pick the right language, since Chinese & English, English, Japanese, Korean, French and German are separate choices and you can only use one per run. Tables, columns and footnotes are read as lines of text, so layout is not preserved and you may need to tidy the order.

Processing happens on your device, one page at a time, so long documents take a while, and how long depends on your hardware and the page count. Keep the tab open until it finishes. The first run downloads the model, about 15 MB, which is then cached. Proofread names, numbers and anything important.

Frequently asked questions

How do I extract text from a scanned PDF?

Drop the PDF onto the page, select the document's language, and click Start Recognition. The text appears page by page, and you can copy it or download a TXT file.

Does this create a searchable PDF?

No. It outputs plain text only and doesn't add a text layer to your PDF. Your original file is left unchanged.

Is my scanned PDF uploaded to a server?

No. The pages are rendered and recognized in your browser. Only the text-recognition model is downloaded, about 15 MB the first time, and the PDF stays on your device.

Which languages work for PDF OCR?

Chinese & English (including traditional Chinese), English, Japanese, Korean, French and German, one per run.

Why is my OCR result inaccurate?

Usually the scan: low resolution, skewed pages, shadows or the wrong language. Use a clear scan of about 300 dpi, make sure the language matches, and proofread the output.

Guides