Generate complete waveform images with audiowaveform
An 800-pixel waveform at 100 pixels per second shows only the first eight seconds of audio. For a complete overview, generate waveform data and fit all its points into the image. This walkthrough produces an 800×200 PNG with separate stereo channels, including sound in the recording’s final partial data window.
Install the tested tools
These commands use Bash on Debian 13, amd64, with audiowaveform 1.10.2 and Debian’s Python 3 package.
You need cURL, CA certificates, and permission to install packages with sudo. The small Python
renderer uses only the standard library.
Run this installation block in a writable directory. It creates audiowaveform-install for the
package download and leaves your shell in its original directory. If that folder already exists,
installation stops; choose a new folder name to repeat it. The pinned package comes from the
official 1.10.2 release.
(
mkdir audiowaveform-install &&
cd audiowaveform-install &&
curl -fsSLo audiowaveform.deb \
https://github.com/bbc/audiowaveform/releases/download/1.10.2/audiowaveform_1.10.2-1-13_amd64.deb &&
printf '%s\n' '4208706c6ae5ffb5761dddf8294eaf648c7ba9fb4d6fc6ec9f021d1bb109ea6e audiowaveform.deb' | sha256sum --check - &&
sudo apt-get update &&
sudo apt-get install -y ./audiowaveform.deb python3 &&
audiowaveform --version &&
python3 --version
)
Expect AudioWaveform v1.10.2. This package targets Debian 13 on amd64; it is not an Ubuntu PPA
installation. See the project’s installation guide
for other systems.
Generate reusable waveform data
Put a WAV recording named input.wav in your working directory. Use that same directory for the
remaining files and commands. The example was tested with 16-bit PCM and floating-point WAV audio.
The tool also documents MP3, FLAC, Ogg Vorbis, and Opus inputs; see its
input-format options
for format and dependency details.
audiowaveform -i input.wav -o waveform.json --zoom 256 --bits 16 --split-channels
Each point in waveform.json stores minimum and maximum amplitudes for a group of 256 samples per
channel. The last group can be shorter. --bits 16 controls the precision of these stored amplitudes,
and --split-channels retains separate channels instead of combining them. For a mono recording,
there is just one channel.
The JSON format
stores drawing data, not playable audio. Keep input.wav if you want to play the
recording or generate more detailed data later. These commands replace existing JSON and PNG files
with the same output names without prompting. Use separate names for outputs you want to keep, and
run the commands sequentially.
Fit every data point into the PNG
Save this as render-waveform.py beside waveform.json:
import json
import subprocess
from pathlib import Path
waveform = json.loads(Path("waveform.json").read_text(encoding="utf-8"))
width = 800
height = 200
points_per_pixel = max(1, (waveform["length"] + width - 1) // width)
zoom = waveform["samples_per_pixel"] * points_per_pixel
subprocess.run(
[
"audiowaveform", "-i", "waveform.json", "-o", "waveform.png",
"--zoom", str(zoom), "--width", str(width), "--height", str(height),
"--no-axis-labels",
"--background-color", "ffffff",
"--waveform-color", "1a73e8,dc2626",
],
check=True,
)
Then render the image:
python3 render-waveform.py
Open waveform.png. A stereo recording has two stacked bands, blue and red, within the total
800×200 image. A mono recording has one blue band. --no-axis-labels removes the labels and border.
The calculation rounds up to a whole number of cached data points per pixel. For example,
2,344 points need three points per pixel to fit within 800 pixels. Multiplying by the cache’s
256 samples per point gives --zoom 768. This includes the last, possibly incomplete group and can
leave unused width. A recording with fewer than 800 cached points leaves more unused width;
generate its JSON with a smaller --zoom, at least two, if you need finer detail.
Why calculate the scale? In version 1.10.2, automatic fitting from cached data
truncates its duration to whole seconds.
The time-to-scale calculation
also rounds samples per pixel down. Consequently, --zoom auto can omit the very end even when
rendering directly from audio. Counting and grouping the cached points avoids both rounding problems.
Choose the interval and appearance
For a deliberate close-up, render from the original audio with a fixed time scale. This command
shows the first eight seconds of input.wav at common sample rates such as 44,100 or 48,000 Hz:
audiowaveform -i input.wav -o waveform-first-8s.png \
--start 0 --pixels-per-second 100 --width 800 --height 200 \
--split-channels --no-axis-labels \
--background-color ffffff --waveform-color 1a73e8,dc2626
Audio after that interval is outside the image. For other sample rates, the interval is approximate because audiowaveform uses an integer number of samples per pixel. A flat close-up does not prove the whole recording is silent: compare it with the complete overview.
In the Python renderer, change width and height to suit the space in your application. The zoom
calculation follows the width. Colors are hexadecimal RGB values without a #; the comma-separated
waveform colors correspond to the channels. Keeping --split-channels when generating the JSON is
what preserves those channels for later rendering. The
image options
describe color and amplitude controls.
Check unexpected results
- Missing or unreadable input: confirm the filename and format, and regenerate the JSON only after fixing the input. A failed generation can leave an older JSON file in place.
- An old image after a failed run: check the command’s exit status and diagnostic. The Python
renderer reports a failed audiowaveform invocation through
check=True; an existing PNG is not proof that the latest render succeeded. - An invalid zoom error: a cached waveform cannot supply more detail than it contains. Generate new data from the audio at a smaller samples-per-point value. The complete-overview calculation above never requests a scale finer than its input cache.
- A missing tail: use the calculated overview scale. Increasing image height or amplitude does not expand the time interval.
Use the waveform in an application
Use the PNG as a static preview beside an audio player, with alternative text appropriate to what the image communicates. Keep the JSON if you want to redraw the same recording with different colors or dimensions without decoding the audio again. Playback and seeking controls belong to your application’s audio player; the PNG itself is a static image.
