Accelerating compression with tar and pigz
Pipe an uncompressed tar archive into pigz to compress a project using several CPU cores. You get
an ordinary .tar.gz file that recipients can extract with gzip and tar, without installing pigz.
The walkthrough below creates that archive, excludes logs, and checks the restored file against
its original.
Check the tools
Use Bash and GNU tar on Linux. The examples were tested with GNU tar 1.34, pigz 2.6, and Bash 5.1 on Ubuntu 22.04, and with GNU tar 1.35, pigz 2.8, and Bash 5.3 on Ubuntu 26.04. On Ubuntu or Debian, install pigz with:
sudo apt-get update && sudo apt-get install pigz
Check the installed versions before continuing:
bash --version && tar --version && pigz --version && gzip --version
The commands below target those GNU/Linux environments. macOS ships a different tar implementation, and package availability on other Linux distributions varies; their setup is outside this walkthrough.
Create a small project
Run the following blocks in the same Bash session. Start in a directory where you can create a new
pigz-demo directory:
mkdir pigz-demo &&
cd pigz-demo &&
mkdir project &&
printf 'Keep this project file.\n' > project/README.txt &&
printf 'Omit this debug log.\n' > project/debug.log
The && chain stops if setup fails. If pigz-demo already exists, choose a different location;
this block deliberately refuses to reuse it. You should now be inside pigz-demo, with two files
under project/.
Compress the project
Run this complete block, including the parentheses:
(
set -e
set -o pipefail
set -o noclobber
tar --exclude='*.log' -cf - project | pigz -p 4 -6 > project.tar.gz
printf 'Created project.tar.gz\n'
)
Here, tar -c collects the directory, and -f - writes the archive to standard output. Pigz
compresses that stream with up to four compression threads at level six, its default level. Keep
project.tar.gz outside project/ so the archive cannot include its own output. Do not add tar’s
-z flag here: that would compress the stream before it reaches pigz.
The output policy is to refuse an existing archive. Bash’s noclobber option prevents the
redirection from replacing an existing regular file, including during a noninteractive rerun.
Choose a new output name to keep another archive. The parentheses limit these shell options to
this operation.
With pipefail, a failure in either tar or pigz makes the pipeline fail; set -e then stops the
block before its success message. Otherwise, pigz could successfully compress an incomplete stream
from a failed tar command. A failed run can leave a partial project.tar.gz; do not use it, even
if it passes a gzip integrity check. Investigate the error and choose a fresh output name before
retrying. See Bash’s pipeline and shell option documentation.
Choose what to leave out
The quoted --exclude='*.log' pattern excludes logs beneath the project directory, including
project/debug.log. Quoting lets tar interpret the wildcard instead of the shell. Remove that
option if logs belong in your archive. For more exclusions, GNU tar accepts repeated --exclude
options or --exclude-from with one pattern per line in a file. These are shell-style patterns,
not regular expressions; see the GNU tar exclusion documentation.
Verify and restore the archive
After a successful creation, check the gzip stream and list its archive members:
gzip -t project.tar.gz && tar -tzf project.tar.gz
A successful gzip -t is silent. For the sample project, tar then prints:
project/
project/README.txt
These checks answer different questions: gzip tests the compressed data’s integrity, while the listing lets you check which paths were archived. Neither proves that every intended source file was captured. The gzip manual documents its integrity test.
Extract into a new directory and compare the restored file with the original:
mkdir restored &&
tar -xzf project.tar.gz -C restored &&
cmp project/README.txt restored/project/README.txt &&
printf 'Restored README.txt matches the original.\n'
The final message appears only if extraction and the byte comparison succeed. The initial mkdir
refuses an existing restored directory, so rerunning this block cannot overwrite a previous
restore. If extraction fails, inspect the error and use a new destination for the next attempt.
For a real project, compare the files you need to recover, not just the sample README.
Choose a thread count and compression level
Start with -6. Try -1 when compression time matters more than size, or -9 when you can spend
more CPU time trying to reduce the output. Increasing the level does not mean faster compression.
Without -p, pigz defaults to the number of online processors; an explicit limit is useful on a
machine doing other work. These settings are described in the pigz manual.
The tiny sample above checks correctness, not speed. Once you have a representative, unchanged project, compare one and four compression threads at the same level:
(
set -e
set -o pipefail
for threads in 1 4; do
printf 'Compression threads: %s\n' "$threads"
time tar --exclude='*.log' -cf - project | pigz -p "$threads" -6 > /dev/null
done
)
Bash’s time measures the whole pipeline; compare the real elapsed times. This discards the
compressed output, so it measures reading and compression without writing an archive to disk.
Repeat on your own data: filesystem caching, available CPUs, storage throughput, and already
compressed files can change the result. There is no fixed speedup to expect.
Pigz parallelizes compression, but ordinary gzip decompression still uses one decompression thread, with helper threads for reading, writing, and checksums. Do not expect the same scaling when extracting. The pigz manual explains this distinction.
Use the archive in a backup job
Before archiving a real project, stop processes that change its files or archive a filesystem snapshot. Tar does not make a live directory into a consistent point-in-time backup. For a scheduled job, use absolute paths, give each run a new output name, and preserve the failure checks before transferring or splitting the result.
This workflow creates one complete local archive. Incremental backup chains, SSH transfers, and split archives need their own failure handling and restore procedures. Keep the last known-good backup until you have checked the new archive and the files you need to recover.
