Duplicate files: find and remove them without losing data
A disk with a few years of work on it almost always holds the same file in three, five, ten places: the copy downloaded from email, the one saved in the case folder, the one attached to a draft, the one in the backup of the backup. Finding them is easy. Removing them without losing anything — and being able to prove afterwards what was removed and where it went — is the real work.
The same name is not the same file
The first instinct is to look for duplicates by name. It is the wrong criterion, and it fails in both directions. Two files called report.txt in two different case folders share a name and have nothing else in common. Conversely, IMG_2041.jpg, IMG_2041 (1).jpg and client photo.jpg may be three bit-for-bit copies of the same picture.
Dates do not help either. Copying a file often resets its creation date, so two identical copies can carry dates years apart, while two different files can share a timestamp down to the second. Names and dates are labels the filesystem attaches to content; the content is something else.
The criterion that does not fail is a fingerprint of the content. Probatio treats two files as identical when they have the same SHA-256 (for a refresher on what a digest guarantees, see the guide on cryptographic hashes). It is worth being precise: this is equality of digests, not a byte-by-byte comparison. Two different files with the same SHA-256 would require a collision nobody knows how to produce, so in practice the two are the same thing; but the certainty you rely on is a cryptographic one, and you should know that.
Why not simply hash everything
The naive method is to compute the SHA-256 of every file and group equal digests. It works, but it costs as much as reading the whole disk: on a two-terabyte drive that is hours, and almost all of that time is spent discovering that obviously different files are, indeed, different.
The Deduplication module works as a funnel instead. Each step discards what the previous one has already proven different, and only the survivors move on to the next, more expensive step.
| Step | What it reads | What it discards |
|---|---|---|
| 1. Walk and group by size | Filesystem metadata only | Files of unique size: they cannot have a twin |
| 2. Digest of the first 64 KiB | 64 KiB per file | Files that differ at the start |
| 3. Digest of the last 64 KiB | One seek and 64 KiB, only on files larger than 64 KiB | Files with the same head but a different tail |
| 4. Full SHA-256 | The whole file | Nothing: it forms the groups |
The order of the steps is not an optional optimisation: it is what makes the search usable on a whole disk rather than a single folder. The first step does not even open a file, because size comes from metadata, and it usually removes the vast majority of candidates. The second and third read at most 128 KiB per file, however large the file is. The full read only touches files that have already passed three tests.
Why the tail as well
A head check on its own fails exactly where it would matter most. Two clips from the same camera, of the same length, can have identical sizes and an identical container header: with the head alone both would go on to the full read, and gigabytes would be read to find out they differ. The tail, on the other hand, almost always differs, and checking it costs one seek and 64 KiB. For files that fit entirely within the first 64 KiB the step is skipped, since it would look at the same bytes again.
The filters cannot lose a duplicate
A fair question: what if a quick filter wrongly discarded two identical files? It cannot happen, by construction. Two identical files necessarily have the same size, the same first 64 KiB and the same last 64 KiB: they pass every filter and reach the final comparison together. The filters only discard files that differ in bytes already read — files that are certainly different.
Two exclusions are on by default and can be turned off: empty files (they would all be «identical», often by the hundred) and hidden files. One is fixed: symbolic links are never followed, because following them would count the same content twice and allow it to be removed through a path that is not the real one. The reading mode is also decided on evidence: before the expensive pass Probatio makes twelve test reads on the largest file and, if the median latency exceeds 2 ms, treats the medium as a spinning disk and reads one file at a time, because there parallel reads slow things down instead of speeding them up. The choice, and the reason, go into the report. If some folders could not be opened, the count appears before the results: «no duplicates» only holds for what was actually read.
Hard links are not copies
A file can have more than one name. A hard link is a second path pointing to the same content on disk: two directory entries, one inode, one allocation. They obviously share a SHA-256, but they are not two copies: deleting one frees not a single byte, because the content is still reachable through the other name.
Probatio recognises them by volume and by the number the filesystem assigns to the content — the inode on macOS and Linux, the file index on Windows, read by opening only the files that appear in a group — flags the group as «hard links» and computes reclaimable space by counting distinct contents only: size times (distinct contents minus one). A group of three 2 GB paths, two of which are the same inode, promises 2 GB of reclaimable space, not 4. Promising space that will never be freed is a subtle way of reporting a false figure.
The only feature that deletes
Everything else in Probatio reads. This feature removes the user's files, and its safeguards live in the engine, not only in the interface: the interface mirrors them, but even a badly built plan is stopped downstream.
First safeguard: the digest is recomputed at the moment of action
Minutes, sometimes hours, pass between the scan and the click on «Run». Meanwhile a file may have been reopened and saved, updated by a sync service, replaced. This is the classic time of check to time of use problem: a check is true when you make it, not necessarily when you act on it. If the program trusted the scan, it would remove a file that is no longer the duplicate you saw.
So before moving or deleting any file, Probatio recomputes its SHA-256 by reading it in full and compares it with the group's digest. If it matches, it proceeds. If it does not, the file is left alone and the result states why: the content changed after the scan, so it was not removed. If the file cannot be read again, it is left alone too. The full re-read costs time, and that is deliberate: it is the point at which this feature stops being dangerous.
Second safeguard: a group is never left empty
In the interface the last copy in a group cannot be marked: at least one file must remain for every content. But the real guarantee is in the engine. If it receives a group with nothing to keep, it refuses it entirely, not halfway: deleting part of it and realising afterwards that it should not have happened would be the worst possible outcome. A path listed both as «keep» and as «remove» is refused as well.
Having a copy to keep in the plan is not enough, though: it must still be on disk, with the same content. Before removing any file of a group, Probatio reads the copy to be kept in full and recomputes its SHA-256. If it has been moved, deleted or modified since the scan, removing the others could leave you with no copy of that content at all: the whole group is refused, and the reason given is that the copy to keep is gone or changed after the scan.
Third safeguard: the default action does not delete
By default files go to the system trash, where the user can recover them with tools they already know. The more cautious alternative is quarantine: a folder you choose, into which each file is moved while preserving the structure of its original path. /Volumes/Work/cases/2023/report.txt ends up in <quarantine>/Volumes/Work/cases/2023/report.txt. Flattening the names would let two report.txt files from different folders overwrite each other, which is exactly the typical deduplication case. If the quarantine folder is on another volume and a direct move fails, the file is copied, and the original is deleted only if the copy succeeded.
Permanent deletion exists, but it has to be chosen and confirmed separately, with a dedicated checkbox.
For disks with thousands of groups there is an automatic selection that keeps one copy per group according to a rule: the oldest, the newest, or the one with the shortest path. The dates used are filesystem dates (creation, or modification when creation is missing), so «the oldest» is usually the original, but not always: restoring from a backup resets dates. Automatic selection prepares the plan; it does not run it. Reviewing it before pressing «Run» is still your job.
The report: what was removed and where it went
The PDF report keeps two moments apart: when the scan was run and, if an action was taken, when the action was run, with the time zone. Hours may pass between the two, and they are two different facts. It then states which folders were searched, whether subfolders were included, how files were read and by which method (size, first and last 64 KiB, full SHA-256); how many files were examined, how many groups were found and how much space is reclaimable. It lists the groups with their SHA-256 and the path of every copy as the scan found them, before any action, and the paths that could not be read, with the total before the list.
If an action was taken, it adds the treatment (trash, quarantine or deletion) and a table giving, for every file removed, its path, size, the SHA-256 re-checked just before removal, and its destination: the quarantine path, «system trash», or «permanently deleted». Then comes the list of files left untouched, with the reason. The operation is also recorded in the app's activity log, with the number of files removed and left untouched and the bytes freed.
Limits you should know
- Identical means same SHA-256, not a byte-by-byte comparison (see above).
- The checks at the moment of action cost full reads: the copy to keep and every file to be removed are read again in full. On groups of large files, running the plan takes as long as reading those files.
- The resume point is written to the disk being examined: to resume an interrupted scan, Probatio saves a
.probatio-dedup.jsonfile in the first folder chosen and removes it when the work is done. On a read-only volume the save fails without stopping the scan. Deduplication is a tool for working copies, not for the original medium of an exhibit. - The report has boundaries: it shows at most 300 groups, and says so when it truncates; for the trash it records the treatment but not the location inside the trash. If you need to know exactly where each file went, quarantine is the right choice.
Why the report matters in forensic work
In forensic work a destructive operation is not forbidden in itself: sometimes it is necessary, for instance to reduce an archive full of duplicates to a manageable working copy. It becomes unacceptable when it cannot be reconstructed. Whoever reads the report must be able to know what was there, what was removed, by which criterion, and where the removed material is now. If any one of those answers is missing, the operation becomes a blind spot, and a blind spot in a chain of work is paid for at the first cross-examination.
That is why Probatio's Deduplication is built the opposite way to a disk cleaner: it moves instead of deleting, it stops when the data is no longer what was verified, it never empties a group, and it writes everything down. The space saved is the result. The report is what makes it defensible.