tar is older than most of the people who type it, and the format underneath has barely changed since the tape drives it was named after. I had used it for years without ever looking inside, so I wrote one by hand from the POSIX specification using nothing but Python's struct module, and checked every claim below against GNU tar 1.35 and Python's own tarfile.
The format turns out to be simple enough to write in forty lines and odd enough that most of its consequences are not what I would have guessed. The oddest one is that tar carefully checksums every header and does nothing at all to protect your data.
Forty lines and a tape drive
A tar archive is a sequence of 512-byte blocks. Each file gets one block of header, followed by its contents padded out to the next multiple of 512, and the archive ends with two blocks of zeros. That is the whole structure. There is no magic number at the start of the file, no table of contents, and nothing at the end except those zeros.
The header has fixed offsets, and almost every number in it is written as octal digits in ASCII text rather than as a binary integer:
put(124, 12, b"%011o\0" % size) # file size: eleven octal digits, as text
put(136, 12, b"%011o\0" % mtime) # modification time: same again
put(148, 8, b" " * 8) # checksum field holds spaces while summing
chksum = sum(h) & 0o777777 # add up all 512 bytes
put(148, 8, b"%06o\0 " % chksum) # then write the total back, in octal
Two files built that way, plus the two zero blocks, came to 3,072 bytes: six blocks of 512. GNU tar listed it and extracted both files correctly, from an archive that no copy of tar had ever touched:
-rw-r--r-- 0/0 11 1970-01-01 02:00 hello.txt
-rw-r--r-- 0/0 12 1970-01-01 02:00 world.txt
The checksum guards the header and nothing else
That checksum is a plain sum of the 512 header bytes, not a CRC and certainly not a hash. More importantly, it only covers the header.
I changed one byte inside the contents of hello.txt, turning first file into Xirst file, and extracted the archive:
exit code: 0 extracted hello.txt -> 'Xirst file'
No warning, no error, exit status zero, and a corrupted file on disk. Then I changed a single byte of the header instead, the first letter of the file name:
tar: This does not look like a tar archive
tar: Skipping to next header
tar: Exiting with failure status due to previous errors
So tar refuses an archive whose file name has been damaged, and cheerfully writes out an archive whose file contents have been damaged. Once you know what the checksum is for, the asymmetry makes sense. It exists so a tape drive could tell a header block from a data block and resynchronise after a bad read, not so you could trust what came out.
The practical consequence is that a bare .tar gives you no integrity guarantee whatsoever. If you compress it, gzip and xz both carry their own checks over the whole stream, so a .tar.gz is protected by the compression rather than by tar, and that is a dependency worth knowing about if you ever store uncompressed archives and assume otherwise.
Eleven octal digits
Writing numbers as text has a hard ceiling built into it. The size field holds eleven octal digits, and the largest eleven-digit octal number is 8,589,934,591, which is one byte short of 8 GiB.
A sparse 9 GiB file costs nothing on disk, so I asked GNU tar to archive one in each of its three formats.
The strict POSIX ustar format refuses outright, and the error message states the exact limit:
tar: value 9663676416 out of off_t range 0..8589934591
The GNU format sets the top bit of the size field to signal that the remaining bytes are a binary integer instead of octal text:
size field = b'\x80\x00\x00\x00\x00\x00\x00\x02@\x00\x00\x00'
decoded as big-endian binary = 9663676416
The POSIX pax format does something more interesting. It writes an extra header in front of the file containing plain key=value records, puts the real size there as decimal text with no length limit, and leaves zero in the ordinary size field:
pax records: 19 size=9663676416 | 30 mtime=1791179820.339886948 | ...
Three formats and three different answers to the same 1980s decision, all still in use. When an old tool chokes on a large archive, this is usually why.
There is no index
Because there is no table of contents, the only way to find a file in a tar archive is to start at the beginning and walk forward one header at a time. Each header's size field tells you how far to jump to reach the next one, and that jump is the only navigation the format has.
To see what that costs, I built an archive of 2,000 files of 50 KB each, about 100 MB, and the same 2,000 files as an ordinary uncompressed ZIP. Then I measured how much of each file a reader had to touch to extract only the last member:
| extracting the last of 2,000 files | read | seeks | time |
|---|---|---|---|
| tar, uncompressed | 3.0% | 2,001 | 34 ms |
| zip | 0.2% | 7 | 3 ms |
Python's tarfile did not read all 100 MB, because it used each size field to seek straight over the file data. It did have to visit every one of the 2,000 headers, with a seek per header, and it was about ten times slower than the ZIP reader, which keeps a central directory at the end of the file and jumped straight to the entry it wanted.
Compression takes away even that shortcut, because a gzip stream cannot be seeked. To reach the last file in a .tar.gz you have to decompress everything in front of it:
.tar.gz, streaming read |
read | time |
|---|---|---|
| first file | 0.1% | 0.3 ms |
| last file | 100% | 180 ms |
Which file you ask for decides the cost. Asking for the first is nearly free, and asking for the last costs the whole archive.
Appending, and why GNU tar reads to the end
The lack of an index has an upside that I had never thought about. Appending a file to a tar archive does not require rewriting anything.
tar -r added a file to the 103 MB archive in no measurable time, and the first 100 MB were byte-for-byte identical afterwards. The archive did not even grow. Tar pads archives out to whole 10,240-byte records, and this one ended in twenty empty blocks of padding, so the new member was simply written over the zeros where the end marker used to be.
The same mechanism lets you append an updated version of a file that is already in the archive:
echo "version 1" > config.txt && tar -cf app.tar config.txt
echo "version 2" > config.txt && tar -rf app.tar config.txt
The archive now holds two entries called config.txt, and extracting it gives you version 2, because later entries overwrite earlier ones as tar writes them to disk. Adding --occurrence=1 gives you version 1 instead.
That explains something that had always puzzled me slightly about GNU tar. Asked for a single file from the .tar.gz above, it took 0.22 seconds whether I asked for the first file or the last. It cannot stop when it finds a match, because a later entry with the same name would be a newer version. With --occurrence=1 the first file came back in 0.00 seconds.
Same files, different archive
Tarring the same three files twice, two seconds apart, produced two archives with different hashes:
one.tar two.tar differ: byte 659, line 1
Byte 659 falls inside the modification time field of the second entry's header. The order of the entries was not the order I had created the files in, and not alphabetical either. It was b, a, c, which is whatever order the filesystem happened to return when tar listed the directory. Owner names, group names and access times can all leak in the same way.
For a build system, or anything that compares archives by hash, that is a real problem, and it has a known fix:
tar --sort=name --mtime='2024-01-01 00:00Z' --owner=0 --group=0 --numeric-owner \
--pax-option=exthdr.name=%d/PaxHeaders/%f,delete=atime,delete=ctime \
--format=posix -cf repro.tar -C src .
Run twice with the files touched in between, that produced identical archives byte for byte.
What I got wrong
Three things, and the third would have put a false number into the figure.
My first test of the 8 GiB limit told me that ustar refused the file, which was what I expected, so I very nearly wrote it down. The error was actually GNU features wanted on incompatible archive format, because I had passed --sparse to save disk space and sparse files are themselves a GNU extension. Tar had rejected the flag before it ever looked at the file's size. Without --sparse, and with the output piped into head so tar would stop after writing the header rather than writing 9 GB, I got the real limit and the real error text.
The second was publishing scripts that did not run. The read-cost measurements depended on test archives that I had generated with a throwaway snippet and never saved, so a fresh clone of the repository could reproduce every claim except the ones in the table. I only found out because I cloned it and ran it, and I have now made the same mistake in two different repositories this month.
The third is the one I most want to flag. My first measurement of Python's convenient getmember() API, on a .tar.gz, said that extracting the last file read 200% of the archive, meaning it decompressed the whole thing twice. That went into the first version of the figure. When I reran it from the clean clone, it read 100%.
The difference turned out to be which program had compressed the archive. In both cases the uncompressed bytes were identical, getmember() scanned the entire archive before extracting anything, and it then needed to seek about 50 KB backwards to reach the last file. With an archive written by the gzip command-line tool, that backward seek restarted decompression from the beginning. With one written by Python's own gzip module, it did not. I have not established why, so the figure shows both, and the useful conclusion does not depend on the answer: if you need one file from a compressed tar, use streaming mode and stop as soon as you find it.
What to take from it
Tar was designed for tape, where the only operation available is reading forward, and almost all of its behaviour follows from that. The headers are checksummed so a reader can find its place again after a bad block. Numbers are text because that was portable between machines that disagreed about binary integers. There is no index because a tape cannot jump to the end to read one, and appending is cheap for the same reason.
Most of the practical advice falls out of it. Do not rely on an uncompressed tar to tell you that your data is intact, because it will not. Expect old tools to fail on files over 8 GiB unless the archive uses the GNU or pax extensions. Use a ZIP, or a tar with an external index, if you need to pull individual files out of a large archive often. And if anything compares your archives by hash, pass the reproducibility flags, because two archives of the same files will otherwise differ even when nothing in them has changed.
The scripts, the hand-built archive and the figure are in a small repository, and reproduce.sh reruns every claim above in order.

Top comments (1)
The header-only checksum makes more sense once you remember what it was for. On a tape, a damaged header means you've lost synchronisation with the block structure and everything after it is unrecoverable. Damaged contents cost you one file and the stream stays parseable. So the design isn't an oversight , it's a triage decision that was rational in 1979 and looks strange on a filesystem.
The practical consequence is the one for a runbook: tar is not an integrity mechanism, so the checksum belongs on the archive, not inside it. sha256sum on the .tar.gz is the actual control, and anyone relying on tar to notice corruption is relying on something never built to.
Two additions.
The 8 GiB figure falls out of eleven octal digits , 8,589,934,591 bytes exactly. pax extended headers remove the limit, which is why archives past that size quietly stop being ustar and start being pax, and why some older readers then fail on files they previously handled.
And the "same files tar to different bytes" result is why reproducible builds need the whole flag set rather than one: --sort=name --mtime= --owner=0 --group=0 --numeric-owner, plus gzip -n to drop the timestamp gzip embeds. Miss any one and the hash moves.