About this tool
Explains scanner IDs, OCR text layers, failed redactions and hidden revisions in scanned PDFs, and which clean-up step removes each one.
Scanned Document Metadata Explainer maps the hidden content of a scanned PDF to the structure that holds it — the /Info dictionary and XMP packet, the invisible OCR text layer drawn in text render mode 3, objects left behind by incremental saves, and marks baked into the page image — then shows which clean-up step removes each one. Stripping properties, re-printing to a new PDF, flattening to images and print-and-rescan all remove different things, and only one of them destroys extractable text. Aimed at anyone emailing a scanned contract, ID copy or filing who wants to know what leaves the building with it.
Open Scanned Document Metadata Explainer on AltFTool — it loads instantly in your browser.
Add your input to the workspace.
Adjust the options until the result looks right.
Copy or download the output and put it to work.
Shows exactly what strip, re-print, flatten and print-and-rescan each remove, instead of guessing.
Most metadata advice stops at document properties and misses the searchable OCR text underneath the image.
Warns when your chosen step destroys the text layer and with it search and screen-reader access.
Yes, if OCR was applied. Searchable scans draw the recognised words invisibly behind the page image using text render mode 3, so every word can be selected, copied and indexed by a search engine even though you only see a picture.
Because a filled rectangle or highlight annotation only sits on top of the content. The original characters remain in the content stream, so selecting the area and pasting it, or deleting the annotation, brings the text straight back. Use a redaction function that removes the underlying content, then verify by copy-pasting the area.
Typically the Producer or Creator string naming the device model and firmware, CreationDate and ModDate accurate to the second with a UTC offset, the scan job settings, and — on scan-to-email from a networked multifunction printer — the directory account of the person who ran the job in the Author field.
Most colour laser printers do. The Machine Identification Code is a faint yellow dot grid encoding the printer's serial number and the date and time of printing, and a high-resolution scan of that page reproduces it. Monochrome laser and inkjet output generally does not carry it.