# Optimize documents

Robot: `/document/optimize`

🤖/document/optimize reduces the file size of PDF documents.

This Robot reduces PDF file sizes. It recompresses images, subsets fonts, and applies various optimizations to reduce file size while maintaining acceptable quality.

## Quality Presets

The Robot supports four quality presets that control the trade-off between file size and quality:

|Preset|DPI|Use Case|Typical Savings|
|-|-|-|-|
|`screen`|72|Screen viewing, smallest files|\~86%|
|`ebook`|150|Good balance of quality/size|\~71%|
|`printer`|300|Print quality|Moderate|
|`prepress`|Highest|Press-ready, largest files|Minimal|

Stage: beta

## Usage example

Optimize PDF file size.

```json
{
  "steps": {
    "optimized": {
      "robot": "/document/optimize",
      "use": ":original",
      "preset": "ebook"
    }
  }
}
```

## Parameters

* `interpolate`: Controls whether Assembly Variables are interpolated for individual instruction fields.

  By default, most Robot instruction fields interpolate Assembly Variables. Set this to `false` to treat every instruction field as literal text, or set an individual field path to `false` to treat only that field as literal text. For Robot-specific fields that are literal by default, set this to `true` or set that field path to `true` to opt back into interpolation.

  Use field names such as `path`, or dotted paths such as `ffmpeg.vf` for nested objects.

* `output_meta`: Allows you to specify a set of metadata that is more expensive on CPU power to calculate, and thus is disabled by default to keep your Assemblies processing fast.

  For images, you can add `"has_transparency": true` in this object to extract if the image contains transparent parts and `"dominant_colors": true` to extract an array of hexadecimal color codes from the image.

  For images, you can also add `"blurhash": true` to extract a [BlurHash](https://blurha.sh) string — a compact representation of a placeholder for the image, useful for showing a blurred preview while the full image loads.

  For videos, you can add the `"colorspace": true` parameter to extract the colorspace of the output video.

  For videos, you can also add `"interlaced": true` to detect whether the video is interlaced. This combines the cheap ffprobe `field_order` flag with a bounded `idet` sampling pass over the first frames of the source, exposing `interlaced`, `field_order`, and a diagnostic `interlace_detection` object under `file.meta`. This is computationally expensive and billed accordingly.

  For audio, you can add `"mean_volume": true` to get a single value representing the mean average volume of the audio file.

  You can also set this to `false` to skip metadata extraction and speed up transcoding.

* `user_meta`: Adds custom metadata to each file emitted by this Robot without modifying the file’s contents.

  The values are merged with any existing `user_meta` carried by the input file. If both objects contain the same key, this Robot’s value takes precedence. Assembly Variables are supported, for example `{ "internal_file_id": "${file.id}" }`.

* `result`: Whether the results of this Step should be present in the Assembly Status JSON

* `queue`: Setting the queue to 'batch', manually downgrades the priority of jobs for this step to avoid consuming Priority job slots for jobs that don't need zero queue waiting times

* `force_accept`: Force a Robot to accept a file type it would have ignored.

  By default, Robots ignore files they are not familiar with.
  [🤖/video/encode](/docs/robots/video-encode.md), for
  example, will happily ignore input images.

  With the `force_accept` parameter set to `true`, you can force Robots to accept all files thrown at them.
  This will typically lead to errors and should only be used for debugging or combatting edge cases.

* `ignore_errors`: Ignore errors during specific phases of processing.

  Setting this to `["meta"]` will cause the Robot to ignore errors during metadata extraction.

  Setting this to `["execute"]` will cause the Robot to ignore errors during the main execution phase.

  Setting this to `true` is equivalent to `["meta", "execute"]` and will ignore errors in both phases.

* `use`: Specifies which Step(s) to use as input.

  * You can pick any names for Steps except `":original"` (reserved for user uploads handled by Transloadit)
  * You can provide several Steps as input with arrays:
    ```json
    {
      "use": [
        ":original",
        "encoded",
        "resized"
      ]
    }
    ```
  * You can also tag input Steps with `as` to pass semantic intent to robots:
    ```json
    {
      "use": [
        {
          "name": ":original",
          "as": "image"
        },
        {
          "name": ":original",
          "as": "mask"
        }
      ]
    }
    ```

  > [!Tip]
  > That's likely all you need to know about `use`, but you can view [Advanced use cases](/docs/topics/use-parameter.md).

* `preset`: The quality preset to use for optimization. Each preset provides a different balance between file size and quality:

  * `screen` - Lowest quality, smallest file size. Best for screen viewing only. Images are downsampled to 72 DPI.
  * `ebook` - Good balance of quality and size. Suitable for most purposes. Images are downsampled to 150 DPI.
  * `printer` - High quality suitable for printing. Images are kept at 300 DPI.
  * `prepress` - Highest quality for professional printing. Minimal compression applied.

* `image_dpi`: Target DPI (dots per inch) for embedded images. When specified, this overrides the DPI setting from the preset.

  Higher DPI values result in better image quality but larger file sizes. Lower values produce smaller files but may result in pixelated images when printed.

  Common values:

  * 72 - Screen viewing
  * 150 - eBooks and general documents
  * 300 - Print quality
  * 600 - High-quality print

* `compress_fonts`: Whether to compress embedded fonts. When enabled, fonts are compressed to reduce file size.

* `subset_fonts`: Whether to subset embedded fonts, keeping only the glyphs that are actually used in the document. This can significantly reduce file size for documents that only use a small portion of a font's character set.

* `remove_metadata`: Whether to strip document metadata (title, author, keywords, etc.) from the PDF. This can provide a small reduction in file size and may be useful for privacy.

* `linearize`: Whether to linearize (optimize for Fast Web View) the output PDF. Linearized PDFs can begin displaying in a browser before they are fully downloaded, improving the user experience for web delivery.

* `compatibility`: The PDF version compatibility level. Lower versions have broader compatibility but fewer features. Higher versions support more advanced features but may not open in older PDF readers.

  * `1.4` - Acrobat 5 compatibility, most widely supported
  * `1.5` - Acrobat 6 compatibility
  * `1.6` - Acrobat 7 compatibility
  * `1.7` - Acrobat 8+ compatibility (default)
  * `2.0` - PDF 2.0 standard
