Create and compare tar archives with Zstd and LZ4
Pipe tar’s output through zstd or lz4 to choose the compressor and its settings explicitly.
The walkthrough below creates both archives, restores their files, and measures compression time
and size on the same input. Each archive gets a new name only after compression succeeds; rerunning
a command cannot replace an existing archive.
Check your tools
Use Linux with Bash, GNU tar, GNU coreutils, diff, Python 3.9 or newer, Zstd, and LZ4 1.10.0 or
newer. The commands below were tested with GNU tar 1.35, Bash 5.3.15, Zstd 1.5.7, LZ4 1.10.0, and
Python 3.14.7. Install missing tools through your distribution’s package manager, then check:
tar --version
bash --version
zstd --version
lz4 --version
python3 --version
ln --version
The tar implementation matters. GNU tar supports --zstd,
but has no --lz4 option. bsdtar supports both options when creating archives
and detects those formats when reading them. This walkthrough uses GNU tar and separate compressors;
it is not a macOS/BSD shell recipe, particularly because ln -T below is a GNU option.
Make a reproducible source directory
Run every Bash block from the same starting directory. Parentheses keep cd and shell options
local to each block. Setup refuses to reuse an existing tar-lab directory, so it cannot overwrite
an earlier experiment.
(
set -euo pipefail
mkdir tar-lab
cd -- tar-lab
mkdir project
python3 - <<'PY'
from pathlib import Path
import random
project = Path("project")
with (project / "records.csv").open("xb") as output:
for index in range(1_000_000):
output.write(f"{index:08d},user-{index % 1000:04d},GET,/assets/app.js,200\n".encode())
(project / "binary.bin").write_bytes(random.Random(0).randbytes(8 * 1024 * 1024))
(project / "empty.txt").touch()
(project / "omit.log").write_text("Exclude this log.\n", encoding="utf-8")
PY
)
This fixture combines structured text with deterministic pseudorandom bytes and an empty file. Keep the source still while archiving it. A successful tar command is not a consistent snapshot of files that an application is actively changing.
Create archives without replacing earlier output
Both examples archive project/ and omit names matching *.log, including nested logs. The quoted
pattern reaches tar without shell expansion. Tar’s -c creates the archive; -f - sends it to
standard output. The compressor writes into a fresh temporary directory beside the final archive,
outside the source tree.
pipefail makes a failed tar producer fail the pipeline even if the compressor succeeds.
It does not protect a file opened by shell redirection: > can truncate an old backup before tar
runs. Here, only temporary output is redirected. After the compressor’s integrity test passes,
GNU ln -T gives those
completed bytes their final name. Without -f, it refuses an existing file, directory, or symlink.
The temporary and final names share one filesystem, which hard links require. The exit trap removes
the temporary directory on normal success or failure.
Create the Zstd archive
Use level three and one compression worker. In the
Zstd CLI manual, -T1 means one
compression worker alongside I/O; it is different from --single-thread.
(
set -euo pipefail
cd -- tar-lab
stage=$(mktemp -d .zstd.XXXXXX)
trap 'rm -rf -- "$stage"' EXIT
tar --exclude='*.log' -cf - project | zstd -3 -T1 -c > "$stage/archive"
zstd -t "$stage/archive"
ln -T -- "$stage/archive" project.tar.zst
printf 'Created project.tar.zst\n'
)
Create the LZ4 archive
Use level one and one compression worker here too. LZ4 1.10.0 added multithreaded compression;
its CLI manual documents -T1.
Older LZ4 releases do not support that option.
(
set -euo pipefail
cd -- tar-lab
stage=$(mktemp -d .lz4.XXXXXX)
trap 'rm -rf -- "$stage"' EXIT
tar --exclude='*.log' -cf - project | lz4 -1 -T1 -c > "$stage/archive"
lz4 -t "$stage/archive"
ln -T -- "$stage/archive" project.tar.lz4
printf 'Created project.tar.lz4\n'
)
Success ends with the corresponding Created message. A rerun returns a nonzero status and leaves
the existing archive’s bytes unchanged, even if project/ has disappeared. To create another
archive, choose another final filename; do not remove a prior backup just to make a rerun succeed.
A forced kill or power loss can leave temporary files. These commands do not provide crash durability.
Restore and compare the files
An integrity test checks the compressed stream, but does not establish that it contains the files you intended to save. Restore into a new directory and compare it with the unchanged source:
(
set -euo pipefail
cd -- tar-lab
zstd -t project.tar.zst
mkdir restored-zstd
zstd -dc project.tar.zst | tar -xf - -C restored-zstd
diff -r --exclude='*.log' project restored-zstd/project
printf 'Zstd restore matches the source.\n'
)
(
set -euo pipefail
cd -- tar-lab
lz4 -t project.tar.lz4
mkdir restored-lz4
lz4 -dc project.tar.lz4 | tar -xf - -C restored-lz4
diff -r --exclude='*.log' project restored-lz4/project
printf 'LZ4 restore matches the source.\n'
)
Expect no diff output and the matching success message. The restored trees contain records.csv,
binary.bin, and empty.txt; omit.log is absent. This compares names and file contents, not owners,
permissions, or all filesystem metadata. A restore rerun refuses the existing restore directory.
An extraction failure can leave a partial new tree, so only treat the restore as successful after
the comparison passes. Use these commands for archives you created and trust.
Measure the tradeoff on the same tar stream
Run this block from the starting directory after setup. It creates a temporary uncompressed tar with fixed member order, timestamps, and ownership, then measures each compressor and decoder. There is one warm-up and five measured runs per codec. Every decoded tar must equal the input bytes. Only the benchmark’s private temporary files are reused; it leaves your archives and restores alone.
python3 - <<'PY'
import hashlib
from pathlib import Path
from statistics import median
import subprocess
from tempfile import TemporaryDirectory
from time import perf_counter
def digest(path):
return hashlib.sha256(path.read_bytes()).digest()
def timed(command, source, destination):
with source.open("rb") as stdin, destination.open("wb") as stdout:
started = perf_counter()
subprocess.run(command, stdin=stdin, stdout=stdout, check=True)
return perf_counter() - started
with TemporaryDirectory(prefix="benchmark-", dir="tar-lab") as temporary:
work = Path(temporary)
raw = work / "input.tar"
with raw.open("xb") as output:
subprocess.run([
"tar", "--sort=name", "--mtime=@0", "--owner=0", "--group=0",
"--numeric-owner", "--exclude=*.log", "-cf", "-", "project",
], cwd="tar-lab", stdout=output, check=True)
expected = digest(raw)
raw_size = raw.stat().st_size
print(f"tar bytes: {raw_size}")
print("codec bytes compressed/tar compress_s decompress_s")
for codec, options in [("zstd", ["-3", "-T1"]), ("lz4", ["-1", "-T1"])]:
packed, decoded = work / codec, work / "decoded.tar"
samples = []
for run in range(6):
compress = timed([codec, *options, "-q", "-c"], raw, packed)
decompress = timed([codec, "-q", "-dc"], packed, decoded)
if digest(decoded) != expected:
raise RuntimeError(f"{codec} changed the tar bytes")
if run > 0:
samples.append((compress, decompress))
size = packed.stat().st_size
print(f"{codec} {size} {size / raw_size:.3f} "
f"{median(c for c, d in samples):.4f} {median(d for c, d in samples):.4f}")
PY
The timings include subprocess startup and local file I/O, but exclude tar creation, member extraction, and the SHA-256 comparison. They are warm-cache measurements, not a disk throughput test. The ratio is compressed bytes divided by uncompressed tar bytes; smaller is better.
On September 24, 2026, the fixture produced a 50,401,280-byte tar on Linux x86-64 with an AMD Ryzen AI Max+ 395 CPU. Using the tool versions listed above, the median elapsed times were:
| Compressor settings | Archive bytes | Compressed/tar | Compress | Decompress |
|---|---|---|---|---|
Zstd -3 -T1 | 8,946,393 | 0.178 | 0.0281 s | 0.0124 s |
LZ4 -1 -T1 | 15,830,951 | 0.314 | 0.0519 s | 0.0332 s |
Zstd produced the smaller archive and finished both operations sooner in this run. The repeated CSV structure and one machine’s timings do not establish a universal ranking. Replace the fixture with representative files before choosing a compressor for your workload, and check that the recipient can decode the format. Short runs are sensitive to other activity on the machine.
Use the result in a backup workflow
Keep compression separate from file selection and backup history. Selecting recently modified files
with find does not track deletions. For a restorable chain of changes, follow the
GNU tar incremental backup walkthrough.
Schedule a backup only after testing its failure and restore paths, and retain earlier successful
archives until your retention policy permits removal.
For SSH transfers, copy a completed archive and verify it at the destination. If you split an archive into parts, retain their exact order for reassembly and verify the reassembled file before restoring it. Neither transport nor splitting changes the compressor choice measured here.
