What does OCR mean for your PDF?

OCR looks at the shapes in a scanned image, matches them to characters, and writes an invisible text layer behind the picture. The page looks unchanged; the words become findable.

Accuracy depends on the scan: 300 DPI, straight pages, good contrast and a common typeface give near-perfect results. Photographs at an angle, faint dot-matrix print and handwriting are where it struggles.

It is also what stands between a scanned statement and a spreadsheet: without a text layer, there are no columns to read, only pixels.

What makes OCR accurate, and what makes it useless?

The input decides almost everything. Recognition is pattern matching against letter shapes, so anything that distorts the shapes costs accuracy — and small losses compound, because one wrong character can break a word, a number or a whole table column.

Where it genuinely fails: handwriting, heavily stylised type, text printed over photographs, and pages photographed at an angle in poor light. A phone snapshot of a page is the single most common reason OCR disappoints.

Expect to proofread numbers whatever the quality. The classic confusions — 0 and O, 1 and l, 5 and S, 8 and B — land hardest in exactly the places they matter most, which is account numbers and totals.

  • Good: 300 DPI, flat page, strong contrast, an ordinary serif or sans typeface, one language.
  • Workable: 200 DPI, slight skew, faded print, mixed type sizes.
  • Poor: photographs of pages, dot-matrix or fax print, colored backgrounds, handwriting.