PDF documents

PDF comparison

Two versions of a contract, the copy received and the signed one, the document filed and the one produced: PDF comparison tells you where each change is, page and context, and what kind it is — words, numbers or dates, punctuation, spacing, capitals, the text of a visible signature. It shows the pages side by side with coloured differences, recognises inserted or removed pages and graphic differences such as stamps and handwritten signatures, verifies the digital signatures of both documents and ends with a plain-language verdict. The two PDFs are not modified.

What it does

What the PDF comparison module does

Word by word, with the position

Each change with its page, removed text in red, added text in green and the context, classified by type: words, numbers or dates, punctuation, spacing, capitals, the text of a visible signature.

Pages side by side

Red removed, green added, orange replaced, and blue for graphic differences — stamps, handwritten signatures, images, lines — computed outside the text, so text shifting by a line does not colour the whole page. Inserted or removed pages are recognised.

Plain-language verdict

From “Identical files” to “Substantially different documents”, with the list of what changes and the caveats. Three modes: strict (everything), balanced (without micro-differences), semantic (only what changes the meaning).

Digital signatures and structure

The signatures of both documents go through Probatio's full verification: integrity, AgID/eIDAS chain, timestamp, changes after signing. Images by hash, fonts, annotations, form fields, saved revisions and metadata are compared too.

Optional semantic analysis

It says whether each change alters the meaning or is a rewording, and finds paragraphs with no counterpart. The model (about 480 MB) is downloaded once from Hugging Face, with hash verification; after that everything stays on the computer.

Scans: straightened and read

Scanned or photographed pages are straightened before comparing (corners adjustable by hand) and read with Tesseract 5 OCR, locally. The PDFs stay intact: the transformation goes into the report.

PDF and JSON report

Verdict, changes, side-by-side pages and the hashes of both files, recomputed at generation time; JSON export for whoever needs to process the data.

Step by step

How it works

  1. Pick the two PDFs: from the sidebar, from Explorer with a right click on two PDFs, from “Compare” or from the saved versions in PDF analysis.
  2. Choose the mode (strict, balanced, semantic) and, if needed, OCR and straightening of scans.
  3. Read the verdict, go through the changes on the side-by-side pages and generate the report.
FAQ

Frequently asked questions

How does it differ from the Diff and Compare modules?
Compare tells whether two files are identical from their hashes; Diff shows the differing bytes. PDF comparison reads the content: two PDFs that differ in their bytes can have the same text, and the verdict explains what really changes — words, numbers, graphics, signatures or metadata only.
Are the documents modified or uploaded?
No. The two PDFs are not touched and the comparison runs on the computer, OCR included. Only the one-off download of the semantic model, if you enable it, and the revocation checks on signature certificates, as in the Digital signature module, use the network.
Does it work with scanned documents?
Yes: pages without text are straightened and read with OCR (Italian and English) before comparing. The result depends on scan quality, and the report shows which pages were read by OCR or straightened.