PDF compression guides
Text PDF vs Scanned PDF
Choose a workflow based on the dominant page representation.
Last updated
A text PDF stores characters and drawing instructions, while a scanned PDF usually stores page images and may also include an OCR text layer. Text PDFs are often smaller and remain searchable at any zoom; scans depend more heavily on image resolution and compression. Identify the document type before optimizing it, because image-focused settings can help scans but may be unnecessary for born-digital text.
How can you tell whether a PDF is text-based or scanned?
Try selecting and searching for a visible phrase, then inspect the page at high zoom. Searchable text does not prove the page is born digital because OCR can place an invisible text layer over a scanned image.1, 2, 3
How should you verify the final PDF?
- Keep the original file unchanged and work from a clearly named copy.
- Test selection, search, page images, and file size on several representative pages rather than only page one.1
- Open the exact downloaded file in a second viewer before sending or uploading it.
- Save the destination receipt or delivery evidence when the document matters.
FAQ
Does searchable text prove that a PDF is born digital?
No. OCR can place a searchable text layer over scanned page images, so inspect selection behavior and the underlying page representation.
Which PDF type usually benefits more from image compression?
Scanned or photo-based PDFs often have more image data to reduce, but the acceptable result depends on the smallest meaningful detail.
Sources
- U.S. National Archives — Scanned Images of Textual Recordsgovernment technical guidanceSource reviewed:
- Adobe Acrobat — PDF Optimizer settingsofficial technical documentationSource reviewed:
- PDFStay — Compress PDF workflowPDFStay first-party methodologySource reviewed: