What is OCR PDF text extraction?

Some PDFs show text that can't be copied or extracted correctly. A scanned document stores each page as a picture, so it has no text to copy at all. Other PDFs do store their text, but with a broken or unusual font encoding: the letters look right on the page, yet copying them gives random characters, empty boxes, or nothing. This happens when a font doesn't say which character each of its letter shapes stands for, or says it wrongly.

OCR (optical character recognition) reads the text the way a person does, from the shapes of the letters on the page. It doesn't use the text stored in the PDF, so it gets the right text out of both kinds of files.

Tool description

This tool extracts the text of a PDF with OCR, right in your browser. It draws each page as an image and recognizes the letters on it with the Tesseract OCR engine. Choose the PDF and the language of the document, extract the text, then copy it or download it as a .txt file. The text of each page starts with a page heading, and you can follow the progress page by page.

Example

A two-column report page whose text copies as garbled characters, such as "еКпФщЮФГиНтЮЧьФщ", comes out as:

Page 1

Quarterly Report

The first column starts here. Revenue increased
during the reporting period because the new
warehouse opened in March and shipping times fell
by two days.

Costs stayed flat. The team hired four engineers and
moved the billing system to the new servers without
any downtime for customers.

The second column continues the report. Customer
satisfaction scores rose to 87 percent, and the support
team answered most tickets within one hour.

Next quarter the company plans to open an office in
Lisbon, translate the website into Portuguese, and
launch the mobile app for tablets.

Features

  • Extracts text from scanned PDFs and from PDFs with broken or unusual font encoding
  • Recognizes text in 19 languages, including Arabic, Chinese, Hindi, Japanese, Korean, and Russian
  • Reads pages with several columns one column after the other
  • Reads pages that are scanned sideways or upside down
  • Reads up to four pages at the same time, depending on how many processor cores your device has
  • Shows the text of each page as soon as it is read, and keeps it if you cancel
  • Runs in your browser, so the PDF never leaves your device

How it works

The tool opens the PDF with PDF.js and draws each page upright at 300 dpi, with its images, text, and drawings. The page is turned into a grayscale image, and Tesseract finds the columns and text blocks on it, then recognizes each line. Because the text is read from the image, the fonts and text stored in the PDF don't matter.

When Tesseract has little confidence in the text it read on a page, the page may be upside down, so the tool reads it again turned around and keeps the reading Tesseract is more confident about.

The document language starts as the language of this site if the tool can read it, and as English otherwise. The first time you extract text in a language, the browser downloads the recognition data for it from this site.

Tips

  • Pick the language the document is written in. With the wrong language, letters such as ä, ß, or é are often misread.
  • If you can copy the text from the PDF as it is, the PDF Text Extractor reads the stored text directly, which is faster and keeps the exact characters.
  • The result is plain text: the lines of each page and blank lines between paragraphs are kept, but fonts, bold, and layout are not.
  • Very large pages, such as posters, are drawn at a lower resolution, so small print on them may be missed.
  • A warning shows how many pages had no text, such as blank pages.
  • Password-protected PDFs can't be read. Remove the password first.