Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Zip files are an example of something that works, are used extensively, and have a number of weird design decisions. Some of this just history, other not so much. But good luck replacing them for customers or friends.
They have some cool advantages. Like being able to read directly from the archive without decompressing the whole thing, per file checksum, and skipping compression on parts which won’t benefit. You can also wedge large amounts of arbitrary data into the file without changing it which… is kinda weird and sometimes useful?
Nasty bits are they used 32 bit ints all over the place, so rely on hacks to support larger files. Can’t support true streaming decompression. Implementations can vary quite a lot, as can compatibility. They also have a number of old and weird features that people don’t really use. Like being able to split a zip file into multiple parts, some ancient compression techniques etc.
by shortercode - Why do they use zip at all with all these design flaws? Wouldn't there be better alternatives?by mrtx01
- Note that bsdtar, which, incidentally, is included in modern Windows versions as tar.exe, can extract as many files as are fully present in truncated ZIP files, because it doesn't read the central header (to detect trunctation) until after extracting or listing files.by jasomill
- It's an interesting tale of software spelunking and edge cases, but I was wondering about this:
> Things like the uncompressed size and CRC-32 can be disregarded.
If you compute a running crc32 then you could check for that value in addition to the uncompressed size, giving you 8 bytes of precision. And if you decompress data as you go, you could check for the uncompressed size as well, giving you 12 bytes of precision. Surely that's enough to make sure you always find the real boundary of a file in the zip.
by Dan42 - It seems technically possible to create an append-only ZIP writer that would only add files at the end of a ZIP archive behind the existing central directory, and then write a new central directory including the new files (and excluding the deleted files). That might make it more difficult to irrecoverably damage the archive as most of it contents will likely be preserved after e.g. an abrupt loss of power.
- > Perhaps surprisingly, the ZIP file puts the central directory - which lists the content of the ZIP file - at the end of the file
This is very common for archive files. It lets you easily append a file to the end of the archive (overwriting the directory) followed by the updated directory. If it were at the start, you’d have to rewrite the entire contents of the archive to grow the directory.
by cmovq - I just tested. 7-Zip is already doing exactly what the author's recovery tool does: if I delete the metadata at the end of the file, it still can open and extract files, while complaining about the harmless "Unexepcet end of data". Further, even if I delete the actual content of the last file, it still can show the directory structure, it just can't extract data of the last file.by CrendKing
Thirty years of creating and storing zip files on every major and some minor OSes, filesystems, sketchy transfer protocols, and unreliable media make me think that the huge number of issues this company has with corrupted ZIP files has nothing to do with the format or with its customers' bad practices. I think their software might just produce bad zip files.> recently we had a particularly bad case: a customer sent us an approximately 1 GB project file which was corrupt, for a game already published to Steam, which they'd spent months working on. They told us all their backups were corrupt too... Predictably, WinRAR's repair tool produced a 1 GB ZIP file that contained nothing.by dotancohen