PDFresh guide
PDF Text Layer vs OCR: What Is the Difference?
Learn the difference between an existing PDF text layer and OCR, why scanned images need recognition, and when PDFresh text extraction can help.
Short answer
A text layer is existing character data inside the PDF. OCR is a separate recognition process that creates new text from page images. PDFresh reads existing text layers; it does not run OCR.
Text layer
A PDF with a text layer contains selectable characters behind or alongside the visible page. Text extraction can read that embedded data without trying to understand the page image.
OCR
OCR analyzes an image of text and guesses the letters. It is useful for image-only scans, but accuracy depends on scan quality, language, fonts, handwriting, and page layout.
Where PDFresh fits
Extract PDF Text loads the selected PDF in the browser and asks PDF.js for the existing text content. It can help you confirm whether a document already has usable text.
Tested example
A born-digital PDF returned selectable paragraphs. A scanned receipt image did not return useful text. A PDF with selectable but badly mapped text returned output that looked different from the page.
Page-level checking
Check multiple pages, especially in long files assembled from different sources. Some pages can have a text layer while others are image-only scans.
Limits
PDFresh does not create a searchable PDF, correct OCR mistakes, fix broken character maps, or guarantee layout-preserving extraction.