What is On-Demand Processing?
On-demand processing creates a media variant only after an application or viewer requests it. Unlike advance generation, it avoids producing outputs that may never be used.
How On-Demand Processing works
On-demand media processing treats a requested transformation specification as the recipe for a derived asset. A cache key normally combines the source identity with normalized operations so equivalent requests reuse one result. The first request enters a queue or worker path, then the completed rendition is stored at an edge or origin for later hits. This pattern fits dynamic resizing, format negotiation, and sparse rendition sets, but it moves capacity planning and failure handling into the request path.
A client authenticates and submits files or references together with workflow instructions. The platform validates the request, schedules dependent operations, records state transitions, and exposes results through a response, polling endpoint, or notification.
Platform concepts become reliable only when their lifecycle is explicit. Authentication, idempotency, retries, timeouts, observability, quotas, and terminal states should be designed together rather than added after failures occur.
Key facts
- 1Canonicalizing operation order, defaults, dimensions, and format parameters prevents semantically identical requests from fragmenting the derivative cache.
- 2A cache miss can trigger a thundering herd when many viewers request one new variant; request coalescing or generation locks limit duplicate work.
- 3Signed transformation URLs can constrain permitted operations and source objects, reducing the risk that an attacker turns the endpoint into an open compute service.
When On-Demand Processing matters
Use this approach when request patterns are sparse or unpredictable to reduce encoding and storage costs. The tradeoff is added latency and processing risk on the first request for each variant.
- Running repeatable upload, import, processing, AI, storage, and notification pipelines.
- Tracking long-running media work independently from an application request.
- Applying credentials, quotas, retries, and error policies consistently across integrations.
Working with platform at scale
Guidance that holds across every platform term in this glossary, not just On-Demand Processing.
What you gain
- Reusable workflows separate application intent from processing infrastructure.
- Stable job identifiers and lifecycle events improve observability and recovery.
- Managed queues and workers let products scale without embedding every media tool.
What it costs
- Synchronous responses are simple but keep connections open while long work executes.
- Aggressive retries improve recovery from transient faults but can duplicate work or overload a dependency.
- Higher concurrency reduces queue time until resource contention or a downstream limit becomes the bottleneck.
Answer these before production
- 1Define authentication, authorization, idempotency, retries, and terminal error behavior.
- 2Observe queue time, execution time, callbacks, and partial results with stable identifiers.
- 3Exercise malformed, duplicate, interrupted, and unauthorized requests before launch.
How Transloadit helps with On-Demand Processing
When On-Demand Processing is relevant to your workflow, you can hand the surrounding platform work to Transloadit instead of maintaining the processing stack yourself. Transloadit models file workflows as reusable Assembly Instructions. Upload, import, processing, AI, storage, delivery, status updates, and error handling can be composed without operating the underlying media tools yourself.
Support for a specific codec, container, parameter, or combination can vary by Robot and processing stack. Check the linked documentation for the exact inputs and outputs available for your use case.