A PDF is not simply a screenshot

A PDF describes how a document should be rendered. It can contain page content, fonts, images, annotations, metadata, optional layers, links, attachments and actions.

That flexibility supports maps, language variants, review notes and interactive forms. It also means a visual review alone does not describe every machine-readable part of the file.

Surfaces worth examining

The presence of one of these structures is not automatically suspicious. The important question is what it contains and whether the downstream AI pipeline will process it.

  • Text outside the visible page or styled to blend into the background
  • Optional content that is currently switched off
  • Annotations with hidden or non-viewing states
  • Embedded metadata, files or other content
  • Links, actions and remote references

Extraction results vary

Different PDF tools can produce different text from the same document. One parser may extract an annotation while another ignores it. OCR may recognise text inside an image that a text-only parser cannot see.

This is why “I read the PDF” and “the AI processed the PDF” are not necessarily equivalent. A useful scan result explains which supported surfaces were checked and where coverage was limited.

What the difference means for AI projects

If instruction-like content reaches extracted text, it may become part of the AI’s working context. The practical response is to preserve the original, scan supported surfaces, review findings and check the stated coverage before use.

For supported PDFs, SourceReady can create a separate cleaner copy and scan that copy again. Visible risky text can remain visible in a cleaner copy, so the second scan and human review still matter.

Read the whole result

Ready means the required supported checks completed and found no supported reason to hold the document. Held means a finding or incomplete required check needs review.

Neither label should be read without its coverage. Decision, findings, coverage and limits together describe what the scan established about that file.

Sources and further reading

Primary guidance used to check the technical claims in this article.