PDF documents

PDF analysis

PDF analysis gathers on one page everything a PDF says about itself and everything that can be measured: hashes, Info and XMP metadata, text statistics with language and readability, page-by-page data, the file's structure and security, images, fonts, attachments and saved versions. The text is not what copy-and-paste gives you: it is rebuilt by interpreting each page's content with the real character widths, so lines and words are the ones you see. The whole PDF is read in the isolated process, as in the reader.

What it does

What the PDF analysis module does

Hashes and metadata

MD5, SHA-1 and SHA-256, file-system dates, Info dictionary and XMP packet. Values written as an internal reference (1733 0 R) are followed to the real text.

Text statistics

Characters by class (letters, accented, digits, punctuation, symbols, unreadable), writing systems, words and distinct words, numbers, sentences, lines, paragraphs, URLs and emails, reading time, most frequent words.

Language and readability

Text language (Italian, English, French, German, Spanish) recognised from function words, Gulpease index for Italian and Flesch for English. If the language declared in the file differs from the text's, it says so.

Page by page

Size, rotation, characters, words, lines, paragraphs, images actually drawn (not merely listed among the resources), fonts, annotations and a “probable scan” flag.

Structure and security

Objects and streams, linearization, tagged PDF, declared conformance (PDF/A, PDF/UA, PDF/X), page labels, encryption and denied permissions, bookmarks, layers, form fields, embedded files, JavaScript actions, annotations by type, links and domains.

Images, fonts, attachments, versions

Images with zoomable thumbnails, fonts, attachments and the document's saved versions, each of which can be compared with the current one in PDF comparison.

PDF report

A report with the tables of every section and the image gallery.

Step by step

How it works

  1. Drop a PDF onto the module, or open it with a right click from Explorer.
  2. Probatio reads it in the isolated process and builds the dossier, with the warnings at the top (scans, unreadable characters, a different declared language).
  3. Generate the PDF report, or move on to PDF comparison for the saved versions.
FAQ

Frequently asked questions

Why can the word count differ from another program's?
Probatio rebuilds the text from the page content using the real character widths, Form XObjects included, instead of copying extracted text. Lines and words are the ones drawn; paragraphs and sentences are estimates, and the report says so.
What does “declared language different from the text” mean?
A PDF can declare its own language (/Lang). Probatio compares it with the language recognised in the text: an Italian manual declaring zh-CN points to a different template or source program, a clue about the document's history.