Harnessing process substitution for zero-disk streaming
Use Bash process substitution when a command needs filenames for streams, and a pipeline when one command can read another’s standard output. This walkthrough shows how to detect failures in both, then upload a compressed tar archive to a local HTTP receiver and compare the received files.
Use Linux with Bash 5.2.21 or newer, GNU sort, diff, cmp, GNU tar 1.35, gzip 1.12 or newer,
cURL 8.5.0 or newer, and Node.js 24.15.0 or newer. The examples were also tested with Bash 5.3.15,
gzip 1.14, cURL 8.22.0, and Node.js 26.8.1. Run saved shell scripts with bash, not sh;
the receiver uses Node.js’s native TypeScript support
and an explicit .mts extension. No packages need installing for the receiver.
First, paste this into Bash to create a small dataset. The parentheses keep your current directory
unchanged, and && stops setup if stream-demo already exists or navigation fails.
(
mkdir stream-demo &&
cd stream-demo &&
mkdir input &&
printf 'pear\napple\n' >input/left.txt &&
printf 'apple\npear\n' >input/right.txt &&
printf '\000\377\200binary\n' >input/bytes.bin &&
: >input/empty
)
Open two terminals in stream-demo. Save the following scripts there and run all remaining commands
from that directory.
Understanding process substitution
With <(command), Bash starts a producer asynchronously and passes a filename referring to its
output. On Linux this commonly looks like /dev/fd/63. It is a stream, so consumers cannot assume
they can seek through it like a regular file. The reverse form, >(command), supplies a filename
you can write to as that command’s input. See the
Bash process substitution manual.
The tempting one-liner diff <(sort file1.txt) <(sort file2.txt) has a trap: if both files are
missing, both producers can fail while diff compares two empty streams and returns zero.
pipefail does not collect the statuses of those background producers.
Save this as compare.sh. It keeps the streams open, records each producer’s PID immediately, and
waits for both producers before reporting the comparison’s status:
#!/usr/bin/env bash
if (( $# != 2 )); then
printf 'Usage: bash compare.sh FILE1 FILE2\n' >&2
exit 2
fi
exec {left}< <(LC_ALL=C sort -- "$1")
left_pid=$!
exec {right}< <(LC_ALL=C sort -- "$2")
right_pid=$!
diff "/dev/fd/$left" "/dev/fd/$right"
comparison=$?
exec {left}<&- {right}<&-
wait "$left_pid"; left_status=$?
wait "$right_pid"; right_status=$?
if (( left_status != 0 || right_status != 0 )); then
printf 'Cannot compare: a sort failed.\n' >&2
exit 2
fi
exit "$comparison"
Run bash compare.sh input/left.txt input/right.txt: no output and status zero mean the sorted
contents match. Different contents return one; a failed producer returns two. The script
deliberately collects statuses without set -e, so one failed command cannot bypass the remaining
wait. Bash documents waiting for process substitutions in its
wait reference.
Benefits of zero-disk streaming
The upload below avoids creating an intermediate archive on the sender. It still reads source files,
and the receiver writes the final archive. “Zero-disk” therefore describes the intermediate transfer,
not the entire workflow. Even sort
can spill large inputs to temporary files.
Avoiding a staged archive saves that archive’s disk space and write/read cycle. It does not guarantee a speedup: compression, storage, and the network can each limit throughput. You also give up a seekable copy that cURL could reread after a failure.
Streaming data between commands
For tar and cURL, an ordinary pipe is sufficient: tar writes an archive to stdout and cURL reads
stdin. Start with a receiver that actually accepts the request. Save this as receiver.mts:
import { createWriteStream } from 'node:fs'
import { createServer } from 'node:http'
import { pipeline } from 'node:stream'
const server = createServer((request, response) => {
if (request.method !== 'PUT' || request.url !== '/archive') {
request.resume()
response.writeHead(404).end()
return
}
pipeline(request, createWriteStream('received.tar.gz', { flags: 'wx', mode: 0o600 }), (error) => {
if (error) {
console.error('Upload failed; inspect received.tar.gz before retrying.')
response.destroy()
return
}
response.writeHead(201).end()
})
})
Append the following startup code to the same file. Port zero asks the OS for an available port; an optional numeric argument selects a particular port.
const port = Number(process.argv[2] ?? 0)
if (!Number.isInteger(port) || port < 0 || port > 65535) {
throw new Error('Port must be an integer from 0 to 65535')
}
server.on('error', (error) => {
console.error('Receiver could not listen:', error.message)
process.exitCode = 1
})
server.listen(port, '127.0.0.1', () => {
const address = server.address()
if (!address || typeof address === 'string') throw new Error('Missing listening address')
console.log(`http://127.0.0.1:${address.port}/archive`)
})
In the first terminal, run node receiver.mts and leave it running. Copy its printed URL. The
receiver binds only to loopback, accepts HTTP/1.1 chunked uploads, and streams bytes into
received.tar.gz using Node’s
pipeline.
The wx flag refuses an existing file. A failed upload can leave a partial file, and a storage
error closes the connection. This small local receiver has no authentication, upload-size policy,
or archive validation; keep it private and use one upload at a time.
Next, save this argument check as the start of upload.sh:
#!/usr/bin/env bash
if (( $# != 2 )); then
printf 'Usage: bash upload.sh DIRECTORY URL\n' >&2
exit 2
fi
if [[ $2 != https://* && ! $2 =~ ^http://127[.]0[.]0[.]1:[0-9]+/ ]]; then
printf 'Use HTTPS, or loopback HTTP for this demo.\n' >&2
exit 2
fi
Append the pipeline to upload.sh:
set -o pipefail
if status=$(tar -czf - -C "$1" . |
curl --disable -fsS --http1.1 --noproxy 127.0.0.1 --globoff \
--connect-timeout 5 --max-time 60 --proto '=http,https' \
--header 'Content-Type: application/gzip' --upload-file - \
--output /dev/null --write-out '%{http_code}' --url "$2"
) && [[ $status == 201 ]]; then
printf 'Archive sent; verify the received files.\n'
else
printf 'Upload failed; inspect the receiver before retrying.\n' >&2
exit 1
fi
In the second terminal, run bash upload.sh input 'COPIED_URL', replacing COPIED_URL with the
receiver’s complete URL. Keep the source directory unchanged during the transfer. The receiver’s
archive lives outside input, so tar cannot accidentally include its own output.
Here, tar -czf - produces a gzip-compressed archive, so the content type is
application/gzip.
cURL’s --upload-file - reads stdin and sends an HTTP PUT.
--disable prevents a personal .curlrc from changing the example’s behavior. There are no
redirects or automatic retries. The script requires HTTP 201 from this receiver as well as a
successful pipeline: pipefail
makes a failed tar visible even when cURL successfully sends the bytes it received.
After the success message, paste this into the second terminal to extract into a new directory and compare the dataset byte for byte:
(
mkdir unpacked &&
tar -xzf received.tar.gz -C unpacked &&
cmp input/left.txt unpacked/left.txt &&
cmp input/right.txt unpacked/right.txt &&
cmp input/bytes.bin unpacked/bytes.bin &&
cmp input/empty unpacked/empty &&
printf 'All four files match.\n'
)
This checks ordinary text, non-text binary bytes, and an empty file. An existing unpacked
directory stops extraction instead of overwriting earlier results. Only extract archives you trust.
Stop the receiver with Ctrl+C when finished.
Troubleshooting and best practices
If tar cannot read its input, the script returns nonzero even if the receiver returns 201. That
HTTP response means it stored a complete request body; it cannot tell whether the producer failed
or the archive contains the intended files. Treat the destination as suspect until verified.
An HTTP rejection, disconnect, or cURL timeout also fails the upload. A disconnect after storage but before the response reaches cURL leaves the result uncertain: the file may already be complete. If the receiver cannot start, check its error before uploading; an occupied port is not proof that your receiver is running.
Do not add --retry to a cURL process reading a pipe and expect it to regenerate the archive.
The consumed bytes cannot be rewound, and cURL does not restart tar. Its
retry documentation also warns about
redirected input. To retry this demo, stop the receiver, inspect and move aside or remove its
received.tar.gz, then start the receiver again and rerun bash upload.sh with the newly printed
URL. Each invocation starts a fresh producer and sends the whole archive. Use a fresh extraction
directory for verification; there is no resume support in this workflow.
Advanced compression options
Keep gzip for this walkthrough so the producer, filename, content type, and extraction command agree. Switching compressors changes all four. Compression tuning is a separate decision from streaming: measure your own data before choosing a different format.
Real-world applicability
A real HTTP destination must accept PUT, the request’s streaming framing, and your expected success status. This is not a multipart form upload or an SSH extraction command. Use HTTPS and the destination’s documented authentication for remote transfers; do not put passwords in a command line. The script allows plain HTTP only for the loopback demonstration and does not follow redirects.
If you need replayable bytes, a known content length, or verification before publication, staging an archive can be the better tradeoff. A production receiver also needs its own validation and publication policy before exposing an uploaded object to other readers.
Choose when to stream
Use process substitution for tools that require stream filenames and a pipeline for a single producer and consumer. In either case, account for every producer’s exit status before calling the operation successful. To explore managed compression instead of maintaining this transfer path, see our File Compressing Service.
