deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Hand-built tar archive shows its checksum guards headers, never file contents

A dev.to author reimplemented tar from the POSIX spec and found its checksum protects only headers, leaving file contents unchecked, alongside an 8 GiB size ceiling and no archive index.

Hand-built tar archive shows its checksum guards headers, never file contents

A tar implementation from scratch

A developer writing on dev.to rebuilt the tar archive format by hand from the POSIX specification, using only Python's struct module, and verified the results against GNU tar 1.35 and Python's own tarfile library. The implementation fits in roughly forty lines, and the exercise surfaced several behaviors of the format that are easy to miss after years of everyday use.

The structure is minimal. An archive is a series of 512-byte blocks: each file contributes one header block followed by its contents padded to the next 512-byte boundary, and the archive terminates with two blocks of zeros. There is no magic number, no central directory, and nothing at the end beyond those zero blocks. Numbers in the header, including file size, modification time and the checksum itself, are stored as octal digits written in ASCII text rather than as binary integers.

The checksum stops at the header

The header checksum is a plain arithmetic sum of all 512 header bytes, computed while the checksum field is temporarily filled with spaces, then written back in octal. It is not a CRC and not a hash. More importantly, it covers only the header.

The author demonstrated the consequence directly. Flipping a single byte inside a file's contents produced an archive that GNU tar extracted without complaint, exiting with status zero and writing a damaged file to disk. Flipping one byte in the header, the first character of the file name, made tar reject the archive outright.

According to the write-up, the asymmetry is deliberate: the checksum exists so a tape drive could tell a header block from a data block and resynchronize after a bad read, not so you could trust what came out of the archive.

An 8 GiB ceiling written in octal

Encoding sizes as eleven octal digits caps the format at 8,589,934,591 bytes, one byte under 8 GiB. The author asked GNU tar to archive a sparse 9 GiB file in each of its three formats and got three different answers. Strict POSIX ustar refuses outright, and its error message names the exact limit. The GNU format flips the top bit of the size field to signal that the remaining bytes are a big-endian binary integer instead of octal text. The POSIX pax format takes a third route, writing an extra header containing plain key=value records, placing the real size there as decimal text with no length limit, and leaving zero in the ordinary size field. When an old tool chokes on a large archive, this divergence is usually why, the author notes.

No index, so lookups walk the file

Because there is no table of contents, the only way to find a file is to walk headers from the start, using each size field to jump over the data that follows. In a benchmark of 2,000 files totalling about 100 MB, extracting the last member required touching 3.0 percent of the bytes across 2,001 seeks and 34 ms with Python's tarfile, versus 0.2 percent, 7 seeks and 3 ms for ZIP, which keeps a central directory at the end of the file. Compression removes even that seek shortcut, since a gzip stream cannot be skipped through: reading the last file in a .tar.gz forced a full decompression of everything in front of it, while reading the first was nearly free.

The missing index has an upside the author had not anticipated. Appending with tar -r rewrites nothing; the new entry is written over the archive's trailing padding blocks, and in the test the archive did not even grow. Appending an updated file that shares a name with an existing entry creates two copies, and extraction yields the later one. That behavior also explains why GNU tar reads a compressed archive to the end even after finding a match, rather than stopping early, unless --occurrence=1 is passed.

Why it matters

A bare .tar gives you no integrity guarantee for the files inside it. The protection people associate with .tar.gz comes from gzip or xz, whose compression formats carry their own checks over the whole stream, a dependency worth knowing about if you store uncompressed archives and assume otherwise. The format's age shows in other ways too: the author found that archiving the same files two seconds apart produced archives differing at a byte inside the modification-time field. None of this is broken behavior; it is a tape-era format doing exactly what it was designed to do, and the practical takeaway for developers is knowing precisely where its guarantees end.

  • #tar
  • #file-formats
  • #data-integrity
  • #python
  • #gnu-tar