Looking inside an archive without extracting it
In a forensic examination, an archive arrives as an exhibit. The instinctive move — double-click, drag out, open — is also the first mistake: extracting creates new files with new dates, and the operating system's archive handler tells you things about the archive that are not true. This guide explains what is really inside a ZIP, and how to document it without touching it.
Three reasons not to double-click
Extracting is an act that alters. To look at one file in an archive of three thousand entries, the system handler materialises all of them: three thousand new files, with new dates, on a disk where indexers and antivirus software will open them on their own. If something hostile was inside, it is now on your filesystem.
The handler lies. Windows File Explorer presents a ZIP as a virtual folder and translates every problem into the language of a copy: «cannot copy the file», or «the compressed folder is invalid» — even on archives that fully conform to the specification, which 7-Zip and unzip open without a single complaint.
A file name is not a string. Inside a ZIP it is a sequence of bytes, and different programs read it in different ways. Showing one reading and keeping quiet about the others loses information.
Probatio's Archives module exists for these three problems.
Reading the index, not the contents
In a ZIP every entry appears twice: a local header right before the compressed data, and a record in the central directory, the index at the end of the file, closed by a terminator (EOCD). Conforming readers start from the end, which is why a self-extracting executable works: a program sits in front, but the index is still at the back.
Probatio does the same. Listing an archive of many gigabytes is almost instant, because it touches no compressed byte to learn what is inside — and it decrypts nothing: in a password-protected archive, names, sizes and dates are visible anyway.
The format is recognised from its bytes, not its extension: a .docx is a ZIP, so is a self-extractor, and a renamed file cannot hide what it is. Besides ZIP and its family (docx, xlsx, apk, jar, epub, WACZ files, BagIt bags), the module opens 7z, TAR and compressed TAR, single gzip, bzip2, XZ and Zstandard streams, CAB, ISO 9660, ar and Debian packages, and CFB containers — those of .msi installers, old .doc files and Outlook .msg files. In an MSI the stream names are encoded in a 64-symbol alphabet that, read literally, looks like a row of ideograms: Probatio decodes them.
Before it even lists anything, Probatio computes the SHA-256 of the archive file. It always does, even when opening fails: a report has to say which file it failed on.
A file name is a sequence of bytes
Each entry's record has a flags field, and its bit 11 declares whether the name is written in UTF-8. When the bit is clear, the PKWARE specification prescribes the old DOS code page, CP437. The trouble is that many programs write the name in UTF-8 without setting the bit. 7-Zip and unzip detect UTF-8 on their own and show the right name; Windows, according to the test bench cited below, applies the system's OEM code page instead — CP850 on an Italian system, and on most Western European ones.
Here is one name, with bit 11 clear:
| Reading | Result |
|---|---|
| Raw bytes | 70 65 72 63 68 c3 a9 2e 74 78 74 |
| UTF-8 | perché.txt |
| CP437 (the prescribed one) | perché.txt |
| CP850 (Windows in Italy) | perch├®.txt |
Which one is «the» name has no single answer, and Probatio does not choose for you: the entry's detail pane shows the three readings, the raw bytes in hexadecimal and the state of bit 11, and flags the entry when the name is UTF-8 but does not say so. In the tree it uses the UTF-8 reading when the bytes are valid UTF-8 — as 7-Zip and unzip do — and falls back to CP437 otherwise. If the name is pure ASCII the readings coincide, and the pane says so.
In a report, when you cite a non-ASCII entry name with bit 11 clear, quote the bytes as well. They are the only datum nobody has interpreted.
Two dates for the same entry
A ZIP entry's main date is the DOS time: two-second resolution, local time of the machine that created the archive, and no time zone. Nobody records which local time that was. Many archives, however, also carry an extra field: the NTFS one, holding the date in UTC in 100-nanosecond steps, or the Unix Extended Timestamp, in UTC seconds.
Probatio shows both and writes them differently: the DOS date with no suffix, because calling a zone-less time UTC would be inventing data, and the extra-field date with a Z. When the two do not agree to the minute, the entry is flagged. The one-second gap in the figure above is not enough to flag it: that is the DOS time's two-second step, not a discrepancy.
Interpreting the discrepancy is up to you. A gap of whole hours is what you would expect: it is the source machine's time zone on that date, daylight saving included. Any other gap deserves an explanation before it goes into a report. In other formats, 7z and TAR record the instant in UTC, CAB has no zone just like ZIP, and ISO 9660 carries its offset from GMT, which Probatio shows as it is.
CRC-32 and hashes, with nothing written
ZIP and 7z declare a CRC-32 of the uncompressed data for every entry. When you compute an entry's hashes, or extract it, Probatio recomputes the CRC over the decompressed bytes and compares it with the declared one. It is the same logic as the acquisition hash of an E01 image, applied to every file in the container. Other formats carry no per-entry CRC, and the field stays empty: empty, not zero.
With the same caveat as the E01: CRC-32 is not cryptographic. A mismatch says the content is not what the archive declares; a match says the archive is consistent with itself, not that nobody ever touched it — whoever alters an entry can recompute its CRC.
Binding an entry to a real fingerprint takes cryptographic hashes. Probatio computes MD5, SHA-1 and SHA-256 of a single entry by decompressing it in memory: the file is never written to disk. Alongside comes the content's real type, inferred from the decompressed bytes rather than the name.
«The compressed folder is invalid»
The «Will Windows Explorer open this archive?» panel turns a single-variable test bench — 19 archives, 5 engines, Italian-language Windows 11 — into a diagnosis. The bench is published in the article «La cartella compressa non è valida»: what Windows really does with ZIP archives (in Italian). Its most surprising finding concerns name length: Explorer applies the path limit to the entry name in isolation — the path inside the archive — without counting the folder you will extract into.
Length is the only thing that makes Explorer refuse the archive outright. The other conditions let it open but make extraction fail, nearly always with a code that does not name the cause:
- Methods other than Stored (0) and Deflate (8) — BZip2, LZMA, Zstandard and the rest: extraction fails with
0x80004005. - AES encryption — the same generic code, without even asking for the password: it looks like a broken archive. That is why Probatio shows the method as recorded in the archive, that is 99, not the inner one that 99 hides.
- Characters Windows does not allow (
: * | ? " < >) —0x80070057, with a partial and inconsistent extraction. - UTF-8 names without bit 11 — accented names corrupted according to CP850.
- Symbolic links — silently turned into text files holding the target's name.
The verdict is one of four — no obstacle, partial extraction, will appear empty, will be refused — and next to it is the longest name found, with its length. When a client says «the archive is corrupted», this is the first place to look.
Split archives, archives within archives
Archives split at the byte level — evidence.7z.001, .002 and so on, or .zip.001 — are not archives in their own right: they are the original cut with scissors. Probatio presents them as a single stream, without reassembling them on disk, starting from any piece; the hash shown is that of the reassembled archive. A gap in the numbering is named — «parts 003, 007 are missing» — instead of the usual «archive damaged». But completeness is inferred from the numbering: if the last piece is missing, the sequence just looks shorter.
An entry that is itself an archive can be opened as one, down to three levels. To do so Probatio places a service copy in its own working folder, not next to the exhibit; the three-level limit is the defence against recursive decompression bombs.
Nothing disappears silently
This is the principle behind the whole module. When you extract, every entry that is not written appears in a list that opens by itself, with the reason next to it: the name is too long for the filesystem, the path would have escaped the destination folder (zip slip), the compression ratio exceeds 1000× and you did not confirm, the declared size exceeds the per-file cap. Someone looking at the destination folder would have no other way to know that something is missing.
Likewise, a corrected name — a .. climb removed, an absolute path made relative, a reserved Windows device name — is extracted and annotated, not dropped. A symbolic link becomes a text file declaring its target. An encrypted entry is called encrypted, not dismissed as an error. Every file written carries its SHA-256 and recomputed CRC. And decompression bombs are recognised before anything is decompressed, from the compression ratio and from entries that share the same compressed bytes.
What the module does not do
- RAR does not open. The format is recognised, but the only available implementation derives from RARLAB's UnRAR source, under a licence incompatible with Probatio. It is a licensing constraint, not a technical one, and the app says so instead of implying the file is broken.
- True multi-volume ZIPs (
.z01,.z02… plus.zip) are recognised but not opened: each volume has its own structure and the records carry disk numbers that would need remapping, so concatenating them is not enough. - Multi-cabinet CAB sets: the listing is read, but a file spanning two cabinets is not extracted.
- ISO: only ISO 9660 is read, with Joliet names. UDF and Rock Ridge are not interpreted: a modern image that keeps its content in UDF comes out nearly empty, and Probatio says so.
- Isolation differs between systems. On macOS the reading process runs in an operating-system sandbox — no network, no launching other programs, writes only to the working folder and the chosen destination — with memory and CPU limits. On Windows it is a separate process with a time limit, but with no operating-system confinement and no memory or CPU limits. The guards on paths and bombs live in the code and apply everywhere.
- The Windows diagnosis applies to ZIPs, and the thresholds were measured on one specific build of Windows 11. Another version might behave differently.
- An entry's hashes are computed in memory. For an entry of many gigabytes, extract it and hash it with the Hash module.
In practice
- Record the archive's hash first. Probatio computes it on opening; for a split set, also hash every volume you received.
- Document the listing before extracting, with the module's PDF report: names, sizes, methods, dates and findings, fixed before anything is touched.
- Extract only what you need, and keep the list of skipped entries together with the extracted files.
- Report both dates, and the DOS one without a zone: adding a
Zto it in a report is a mistake. - Before writing «archive damaged», check the Windows panel. Often the archive is intact, and it is Explorer that cannot read it.
An archive is a small filesystem packaged by someone else, with its own dates, its own names and its own unintended lies. Looking at it without extracting it is not excessive caution: it is the only way to describe the exhibit as it was, before your disk becomes part of it.