How to Make a PDF Searchable
A scanned PDF is a photograph of paper — there is no text in the file, only pixels. You cannot search, select, copy, or use a screen reader on it. Making it searchable means running OCR (optical character recognition) to add a text layer that sits behind the image. This guide covers the three kinds of OCR output, the local tools that do this well, and why IXPDF does not do OCR in the browser today.
The three OCR output types
OCR tools can produce three different kinds of output. Which one you want depends on what you will do with the result.
1. Searchable PDF (image + invisible text)
The most common and usually the right choice. The original page image stays visible — layout, signatures, stamps, handwriting all look exactly as they did. Behind the image, OCR adds an invisible text layer with the recognized characters. You can search, select, copy, and use a screen reader; what you see on screen is the original scan. This is what Adobe Acrobat, Tesseract, and ABBYY produce by default.
Best for: archiving scanned documents where the original appearance matters (contracts, receipts, signed letters).
2. Text-only PDF (recognized text, no image)
OCR replaces the image entirely with recognized text. The result is small, fully selectable, and looks like a normal text PDF — but the original layout is approximate, and any content OCR could not recognize (signatures, stamps, handwriting) is gone.
Best for: when you only need the text and the original appearance does not matter (e.g., extracting the body of a scanned letter into a searchable file for a knowledge base).
3. Dual-layer PDF (image + text, both visible)
A less common format that shows both the original image and the recognized text side by side, or overlays the text on the image with the image faded. Used in legal and archival workflows where reviewers need to compare the OCR output against the original to catch misreads. ABBYY FineReader and Adobe Acrobat Pro support this.
Best for: legal review, archival quality control, any workflow where OCR accuracy must be verified against the source.
Local OCR tools
Every tool below runs on your machine. None of them upload your document. That matters: a scanned contract, medical record, or tax form should not be sent to a "free online OCR" service — at that point the content is on someone else's server.
1. Adobe Acrobat Pro (paid, the gold standard)
Open the scanned PDF, then Tools > Scan & OCR > Recognize Text > In This File. Choose the language and output type (Searchable Image is the default and usually right). Acrobat's OCR is the most accurate on real-world scans, handles multi-column layouts well, and lets you correct misrecognized words inline. The downside is cost — Acrobat Pro is a subscription.
2. Tesseract (free, open source, command line)
The most accurate free OCR engine. Install fromthe Tesseract GitHub project or your package manager (brew install tesseract on macOS,apt install tesseract-ocr on Linux). Then:
tesseract input.pdf output -l eng pdf
This produces output.pdf as a searchable PDF with the original image and an invisible English text layer. Tesseract is not as accurate as Acrobat on complex layouts but is excellent on clean single-column scans. It supports 100+ languages — pass-l eng+fra for mixed English/French documents.
3. macOS Preview Live Text (free, macOS 12+)
On macOS Monterey and later, Preview can recognize text in scanned PDFs and images using the system's Live Text feature. Open the PDF in Preview, and you can select, copy, and search text directly — macOS adds the text layer on the fly. To save the searchable version,File > Export as PDF. Live Text is fast and accurate on English documents but limited in language support compared to Tesseract.
4. ABBYY FineReader (paid, high accuracy)
A commercial OCR tool known for high accuracy on difficult scans — handwriting, low contrast, mixed fonts, multi-column layouts. It supports dual-layer output for verification. Use this when Tesseract is not accurate enough and you cannot justify Acrobat Pro's subscription. FineReader is a one-time purchase, not a subscription.
Step-by-step: make a PDF searchable with Tesseract
- Install Tesseract. On macOS,
brew install tesseract. On Linux,apt install tesseract-ocr. On Windows, download the installer from the Tesseract GitHub project. - Install language data. English is included by default. For other languages, install the matching language pack (e.g.,
brew install tesseract-langon macOS). - Open a terminal. Navigate to the folder containing the scanned PDF.
- Run OCR.
tesseract input.pdf output -l eng pdf - Check the result. Open
output.pdf, search for a word you can see in the image, and try to select a paragraph. If search and selection work, the text layer is in place. - Proofread. OCR misreads. Copy a paragraph into a text editor and compare it against the image. Fix any critical misrecognitions in the source document if you will rely on the text.
OCR accuracy: what to expect
OCR accuracy depends almost entirely on the quality of the source image. A clean 300-DPI scan of a printed page in a common font (Arial, Times New Roman) will recognize at 98–99% character accuracy. A 150-DPI scan of a photocopied page with a decorative font may drop to 90% or lower — that is 1 wrong character in 10, enough to break search and copy-paste.
Things that hurt accuracy: low resolution (under 200 DPI), skew (pages scanned at an angle), noise (speckles from a dirty scanner glass), mixed fonts, handwriting, multi-column layouts, tables, and text on colored backgrounds. If your scan has any of these, expect to proofread carefully.
Common mistakes
- Using a free online OCR service for a sensitive document. The file is uploaded to a server. Use a local tool — Tesseract, Acrobat, or macOS Live Text.
- Trusting OCR output without proofreading. OCR misreads. If you are copying text into a contract, a citation, or a database, verify every word against the image.
- Choosing text-only output when you need the original appearance. Text-only OCR discards the image. If signatures, stamps, or layout matter, use searchable PDF output instead.
- OCR-ing a PDF that already has text. If the file already has a text layer (test with Ctrl+F), OCR is unnecessary and may degrade the existing text. SeeWhy Can\u2019t I Select Text in a PDF? to confirm the file actually needs OCR.
- Skipping the language pack. Tesseract defaults to English. OCR-ing a French document with the English language model produces garbage. Always pass
-lwith the correct language.
Edit the metadata of your searchable PDF
Once OCR has added a text layer, use IXPDF's Edit Metadata tool to set the title, author, subject, and keywords — locally in your browser, no upload.
Open Edit MetadataTry it now
Ready to put this into practice?
Run the tool locally in your browser — no upload, no account, no watermarks. Your file never leaves your device.