# Generate waveform images from audio

Robot: `/audio/waveform`

🤖/audio/waveform generates waveform images for your audio files and allows you to change their colors and dimensions.

We recommend that you use an [🤖/audio/encode](/docs/robots/audio-encode.md) <dfn>Step</dfn> prior to your waveform <dfn>Step</dfn> to convert audio files to MP3. This way it is guaranteed that [🤖/audio/waveform](/docs/robots/audio-waveform.md) accepts your audio file and you can also down-sample large audio files and save some money.

Similarly, if you need the output image in a different format, please pipe the result of this <dfn>Robot</dfn> into [🤖/image/resize](/docs/robots/image-resize.md).

Stage: ga

## Usage example

Generate a 400×200 waveform in \`#0099cc\` color from an uploaded audio file:

```json
{
  "steps": {
    "waveformed": {
      "robot": "/audio/waveform",
      "use": ":original",
      "width": 400,
      "height": 200,
      "outer_color": "0099ccff",
      "center_color": "0099ccff"
    }
  }
}
```

## Parameters

* `interpolate`: Controls whether Assembly Variables are interpolated for individual instruction fields.

  By default, most Robot instruction fields interpolate Assembly Variables. Set this to `false` to treat every instruction field as literal text, or set an individual field path to `false` to treat only that field as literal text. For Robot-specific fields that are literal by default, set this to `true` or set that field path to `true` to opt back into interpolation.

  Use field names such as `path`, or dotted paths such as `ffmpeg.vf` for nested objects.

* `output_meta`: Allows you to specify a set of metadata that is more expensive on CPU power to calculate, and thus is disabled by default to keep your Assemblies processing fast.

  For images, you can add `"has_transparency": true` in this object to extract if the image contains transparent parts and `"dominant_colors": true` to extract an array of hexadecimal color codes from the image.

  For images, you can also add `"blurhash": true` to extract a [BlurHash](https://blurha.sh) string — a compact representation of a placeholder for the image, useful for showing a blurred preview while the full image loads.

  For videos, you can add the `"colorspace": true` parameter to extract the colorspace of the output video.

  For videos, you can also add `"interlaced": true` to detect whether the video is interlaced. This combines the cheap ffprobe `field_order` flag with a bounded `idet` sampling pass over the first frames of the source, exposing `interlaced`, `field_order`, and a diagnostic `interlace_detection` object under `file.meta`. This is computationally expensive and billed accordingly.

  For audio, you can add `"mean_volume": true` to get a single value representing the mean average volume of the audio file.

  You can also set this to `false` to skip metadata extraction and speed up transcoding.

* `user_meta`: Adds custom metadata to each file emitted by this Robot without modifying the file’s contents.

  The values are merged with any existing `user_meta` carried by the input file. If both objects contain the same key, this Robot’s value takes precedence. Assembly Variables are supported, for example `{ "internal_file_id": "${file.id}" }`.

* `result`: Whether the results of this Step should be present in the Assembly Status JSON

* `queue`: Setting the queue to 'batch', manually downgrades the priority of jobs for this step to avoid consuming Priority job slots for jobs that don't need zero queue waiting times

* `force_accept`: Force a Robot to accept a file type it would have ignored.

  By default, Robots ignore files they are not familiar with.
  [🤖/video/encode](/docs/robots/video-encode.md), for
  example, will happily ignore input images.

  With the `force_accept` parameter set to `true`, you can force Robots to accept all files thrown at them.
  This will typically lead to errors and should only be used for debugging or combatting edge cases.

* `ignore_errors`: Ignore errors during specific phases of processing.

  Setting this to `["meta"]` will cause the Robot to ignore errors during metadata extraction.

  Setting this to `["execute"]` will cause the Robot to ignore errors during the main execution phase.

  Setting this to `true` is equivalent to `["meta", "execute"]` and will ignore errors in both phases.

* `use`: Specifies which Step(s) to use as input.

  * You can pick any names for Steps except `":original"` (reserved for user uploads handled by Transloadit)
  * You can provide several Steps as input with arrays:
    ```json
    {
      "use": [
        ":original",
        "encoded",
        "resized"
      ]
    }
    ```
  * You can also tag input Steps with `as` to pass semantic intent to robots:
    ```json
    {
      "use": [
        {
          "name": ":original",
          "as": "image"
        },
        {
          "name": ":original",
          "as": "mask"
        }
      ]
    }
    ```

  > [!Tip]
  > That's likely all you need to know about `use`, but you can view [Advanced use cases](/docs/topics/use-parameter.md).

* `ffmpeg`: A parameter object to be passed to FFmpeg. If a preset is used, the options specified are merged on top of the ones from the preset. For available options, see the [FFmpeg documentation](https://ffmpeg.org/ffmpeg-doc.html). Options specified here take precedence over the preset options.

* `ffmpeg_stack`: Selects the FFmpeg stack version to use for encoding. We currently recommend using "v7". The exact versions "v6.0.0", "v7.0.0", and "v8.0.0" are legacy values that remain accepted for backward compatibility. Deprecated "v5.x" values are also accepted.

* `format`: The format of the result file. Can be `"image"` or `"json"`. If `"image"` is supplied, a PNG image will be created, otherwise a JSON file.
  When `style` is `"spectrogram"`, only `"image"` is supported.

* `width`: The width of the resulting image if the format `"image"` was selected.

* `height`: The height of the resulting image if the format `"image"` was selected.

* `antialiasing`: Either a value of `0` or `1`, or `true`/`false`, corresponding to if you want to enable antialiasing to achieve smoother edges in the waveform graph or not.

* `background_color`: The background color of the resulting image in the "rrggbbaa" format (red, green, blue, alpha), if the format `"image"` was selected.

* `center_color`: The color used in the center of the gradient. The format is "rrggbbaa" (red, green, blue, alpha).

* `outer_color`: The color used in the outer parts of the gradient. The format is "rrggbbaa" (red, green, blue, alpha).

* `style`: Waveform style version.

  * `"v0"`: Legacy waveform generation (default).
  * `"v1"`: Advanced waveform generation with additional parameters.
  * `"spectrogram"`: Spectrogram visualization showing frequency content over time.

  For backwards compatibility, numeric values `0` and `1` are also accepted and mapped to `"v0"` and `"v1"`.

* `split_channels`: Available when style is `"v1"`. If set to `true`, outputs multi-channel waveform data or image files, one per channel.

* `zoom`: Available when style is `"v1"`. Zoom level in samples per pixel. This parameter cannot be used together with `pixels_per_second`.

* `pixels_per_second`: Available when style is `"v1"`. Zoom level in pixels per second. This parameter cannot be used together with `zoom`.

* `bits`: Available when style is `"v1"`. Bit depth for waveform data. Can be 8 or 16.

* `start`: Available when style is `"v1"`. Start time in seconds.

* `end`: Available when style is `"v1"`. End time in seconds (0 means end of audio).

* `colors`: Available when style is `"v1"`. Color scheme to use. Can be "audition" or "audacity".

* `border_color`: Available when style is `"v1"`. Border color in "rrggbbaa" format.

* `waveform_style`: Available when style is `"v1"`. Waveform style. Can be "normal" or "bars".

* `bar_width`: Available when style is `"v1"`. Width of bars in pixels when waveform\_style is "bars".

* `bar_gap`: Available when style is `"v1"`. Gap between bars in pixels when waveform\_style is "bars".

* `bar_style`: Available when style is `"v1"`. Bar style when waveform\_style is "bars".

* `axis_label_color`: Available when style is `"v1"`. Color for axis labels in "rrggbbaa" format.

* `no_axis_labels`: Available when style is `"v1"`. If set to `true`, renders waveform image without axis labels.

* `with_axis_labels`: Available when style is `"v1"`. If set to `true`, renders waveform image with axis labels.

* `amplitude_scale`: Available when style is `"v1"`. Amplitude scale factor.

* `compression`: Available when style is `"v1"`. PNG compression level: 0 (none) to 9 (best), or -1 (default). Only applicable when format is "image".

* `color_map`: Available when style is `"spectrogram"`. Color scheme for the spectrogram visualization. Defaults to `"viridis"`.

* `frequency_scale`: Available when style is `"spectrogram"`. Frequency scale for the spectrogram. `"linear"` shows frequencies evenly spaced, `"logarithmic"` emphasizes lower frequencies. Defaults to `"logarithmic"`.

* `frequency_min`: Available when style is `"spectrogram"`. Minimum frequency in Hz to display. Defaults to `0`.

* `frequency_max`: Available when style is `"spectrogram"`. Maximum frequency in Hz to display. Defaults to half the sample rate (Nyquist frequency).

* `legend`: Available when style is `"spectrogram"`. Whether to include a legend showing the frequency and time scales. Defaults to `false`.

* `gain`: Available when style is `"spectrogram"`. Linear gain factor for spectrogram intensity. Defaults to `1`.

* `orientation`: Available when style is `"spectrogram"`. Orientation of the spectrogram. `"horizontal"` shows time on the x-axis (default), `"vertical"` shows time on the y-axis.
