What is VMAF?
VMAF is a full-reference perceptual metric developed by Netflix with academic collaborators to predict subjective video quality from reference and distorted sequences. It combines multiple image-quality features into a score.
How VMAF works
VMAF evaluates an impaired sequence against a temporally and spatially corresponding source, extracts several perceptual features, and fuses them through a trained model. This makes it more viewing-oriented than a single pixel-error formula, but the result still reflects the selected model, preprocessing, and pooling method. Encoding pipelines use it for experiments, regression tests, and ladder analysis, then confirm important decisions with scene inspection and playback evidence.
A demuxer separates tracks from the container, decoders turn compressed streams into frames or samples, and filters apply spatial or temporal changes. Encoders compress the transformed tracks before a muxer writes the chosen output container.
Video compatibility is the product of codec, container, profile, level, frame rate, color, audio, and subtitles. Validate the complete output on target devices because a playable file on one decoder may fail or look different on another.
Key facts
- 1Reference and distorted frames must represent the same moments and comparable image geometry; offsets, dropped frames, or inconsistent scaling can measure misalignment instead of degradation.
- 2A VMAF result is tied to a particular model and tool configuration. Scores produced with different models or preprocessing settings should not be assumed directly comparable.
- 3A single pooled score can conceal a short, severely damaged scene. Per-frame traces and lower-tail or scene-level summaries help expose localized quality failures.
When VMAF matters
Encoding teams calculate VMAF when comparing codecs, settings, transmission variants, or bitrate ladders. Results depend on the selected model and reference, so one score should not replace playback testing.
- Preparing uploaded video for web, mobile, connected-TV, social, or editorial playback.
- Creating clips, thumbnails, captions, alternate aspect ratios, and adaptive renditions.
- Normalizing camera, screen-recording, and user-generated files into predictable outputs.
Working with video at scale
Guidance that holds across every video term in this glossary, not just VMAF.
What you gain
- Standardized derivatives make diverse source files playable on target devices.
- A retained master can feed many resolutions, aspect ratios, codecs, and channels.
- Automated inspection and transformation make large upload volumes consistent.
What it costs
- More efficient codecs can lower bitrate at similar quality but usually cost more compute and may have narrower support.
- Higher resolutions and frame rates preserve more detail and motion while increasing processing and delivery requirements.
- Fast encoding settings improve throughput but can produce larger files or lower quality than slower analysis.
Answer these before production
- 1Inspect codec, container, dimensions, frame rate, color, audio, and subtitle tracks.
- 2Test visual quality and playback support across the slowest and oldest target devices.
- 3Preserve a suitable master before applying lossy, destructive, or delivery-specific changes.
How Transloadit helps with VMAF
When VMAF is relevant to your workflow, you can hand the surrounding video work to Transloadit instead of maintaining the processing stack yourself. Transloadit can transcode, resize, rotate, trim, concatenate, merge, watermark, subtitle, and generate video derivatives, then export each result as part of the same observable workflow.
Support for a specific codec, container, parameter, or combination can vary by Robot and processing stack. Check the linked documentation for the exact inputs and outputs available for your use case.