Extract text and images from PDFs with /document/extract
Your users upload PDFs, but your search index, your RAG pipeline, and your compliance archive need the words inside them. Add one Step to your Assembly to return a PDF's text, per page and as JSON if you like, along with its embedded images. You can feed those results into your existing tools without running a separate extraction service.
The 🤖 /document/extract Robot is generally available.
What you can build with it
- Search. Index each page's text, so a hit can link straight to page 7 of a contract.
- RAG and LLM pipelines. Pages make natural chunks for embeddings, and there is no parsing service for you to run.
- Compliance archives. Keep a plain-text copy next to every original, so retention and discovery tools can read it.
- Auto-filled forms. Prefill fields from the text of an uploaded document. To pull out named fields such as invoice numbers or totals, pass the text to an LLM, as our document intelligence pipeline shows.
Because it is a Robot, it runs in the same
Assembly as the rest of your file handling. The upload
arrives, 🤖 /document/thumbs renders a preview, /document/extract
reads the text, and a storage Robot such as 🤖 /s3/store keeps all three.
Your app gets one Assembly result instead of coordinating separate services.
What it extracts
- Text from the PDF's text layer. This is the default (
text_method: "native"). It uses Poppler'spdftotextin layout mode, so it is fast, involves no OCR fees, and keeps columns and table rows aligned with spaces. Picktext_format"txt"or"json", andtext_granularity"document"(one result) or"page"(one result per page). - Text from scans. With
text_method: "auto", a PDF without a text layer is handed to 🤖 /document/ocr. Useocr_provider: "gcp"andtext_format: "txt"for the recognized text.text_method: "ocr"always uses OCR. - Embedded images. These are the raster images stored inside the PDF, such as photos, scanned
figures, and logos. Identical images are returned once (
dedupe_images), andmin_image_width,min_image_height, andmin_image_byteslet you drop icons and spacers.image_formatpreserves the embedded format where possible or converts to PNG or JPEG.
Every extraction result carries meta.document_extract, which identifies text or images and, for
text, whether it came from the text layer or from OCR. Native per-page text and image results also
identify their page. Document-level metadata, such as the page count, author, and creation date,
is in the upload's meta when available in the PDF.
The Robot returns plain text or JSON, not Markdown, and it does not parse tables into rows and
cells: native extraction returns headings, lists, and tables as layout-preserved text. To render
Markdown to HTML or PDF, use 🤖 /document/convert.
/document/extract does not render pages either; that is the job of
/document/thumbs.
Try it on a real PDF
We will use the Bitcoin whitepaper, a nine-page PDF from our demo inputs. You need Node.js 20.10 or newer and a Transloadit account.
mkdir pdf-text && cd pdf-text
npm init -y && npm pkg set type=module
npm install @transloadit/node
curl -fsSLo bitcoin.pdf https://demos.transloadit.com/inputs/bitcoin.pdf
Put your Auth Key and Secret from the Console in a .env file:
TRANSLOADIT_KEY=MY_AUTH_KEY
TRANSLOADIT_SECRET=MY_SECRET_KEY
These Assembly Instructions extract the text as
JSON and render a preview of page 1 in the same Assembly. Save them as instructions.json:
{
"steps": {
"text": {
"robot": "/document/extract",
"use": ":original",
"extract": ["text"],
"text_format": "json"
},
"preview": {
"robot": "/document/thumbs",
"use": ":original",
"page": 1,
"width": 600,
"format": "jpg"
}
}
}
The script uploads the PDF, waits for the Assembly, downloads the JSON, and prints one line per page:
import { readFile } from 'node:fs/promises'
import { Transloadit } from '@transloadit/node'
const client = new Transloadit({
authKey: process.env.TRANSLOADIT_KEY,
authSecret: process.env.TRANSLOADIT_SECRET,
})
const assembly = await client.createAssembly({
files: { document: './bitcoin.pdf' },
params: JSON.parse(await readFile('./instructions.json', 'utf8')),
waitForCompletion: true,
})
const res = await fetch(assembly.results.text[0].ssl_url)
if (!res.ok) throw new Error(`Could not download the text: HTTP ${res.status}`)
const { pages } = await res.json()
for (const { page, text } of pages) {
const words = text.split(/\s+/).filter(Boolean).length
const firstLine = text.trim().split('\n')[0].replaceAll(/\s+/g, ' ')
console.log(`Page ${page}: ${words} words, starts with "${firstLine.slice(0, 50)}"`)
}
console.log(`Preview of page 1: ${assembly.results.preview[0].ssl_url}`)
$ node --env-file=.env extract.js
Page 1: 461 words, starts with "Bitcoin: A Peer-to-Peer Electronic Cash System"
Page 2: 430 words, starts with "2. Transactions"
Page 3: 526 words, starts with "4. Proof-of-Work"
Page 4: 488 words, starts with "New transaction broadcasts do not necessarily need"
…
Page 9: 159 words, starts with "References"
Preview of page 1: https://…/d650f05c….jpg
Our Node.js 20.10 production test took about 14 seconds for upload and processing; timing varies.
Here is the start of the JSON file the text Step produced, bitcoin-text.json:
{
"text": " Bitcoin: A Peer-to-Peer Electronic Cash System\n\n …",
"pages": [
{
"page": 1,
"text": " Bitcoin: A Peer-to-Peer Electronic Cash System\n\n … Abstract. A purely peer-to-peer version of electronic cash would allow online\n …"
},
{
"page": 2,
"text": "2. Transactions\nWe define an electronic coin as a chain of digital signatures. …"
}
// … pages 3 to 9
]
}
The top-level text holds the whole document, with form-feed characters (\f) separating pages.
pages gives you the same text already split, ready to index or embed per page. Result URLs are
temporary, so add a storage Step when you want to keep the files.
Pull out the images
The whitepaper's diagrams are vector drawings, so asking it for images returns nothing. A PDF exported from Word is a better test. Our demo copy of AWS's Architecting for the Cloud whitepaper has 42 pages, a logo on every one of them, and a single architecture diagram:
{
"steps": {
"images": {
"robot": "/document/extract",
"use": ":original",
"extract": ["images"],
"min_image_width": 300
}
}
}
The Step returns one PNG: the 1,016 × 594 diagram from page 17. The PDF contains 43 raster images
before deduplication, including 42 copies of the logo. Transparency masks are omitted by default
(include_image_masks: false), deduplication collapses the repeated logo, and min_image_width
drops it. Set include_image_masks: true if you also want the separate transparency masks.
Handle scans and other document formats
For a scanned contract without a text layer, set text_method to "auto". The Robot falls back to
OCR only when native extraction finds no text:
{
"steps": {
"text": {
"robot": "/document/extract",
"use": ":original",
"extract": ["text"],
"text_method": "auto",
"text_format": "txt",
"ocr_provider": "gcp"
}
}
}
We ran these Instructions on the whitepaper and on a nine-page scan of it. Including the upload,
the original took about two seconds and was read natively. The scan fell back to OCR and took
about 12 seconds, with its result marked "method": "ocr". Timing varies with the document and
provider.
The Robot reads PDFs only and skips other files. For Word, PowerPoint, or Excel uploads, convert
them first with /document/convert and point /document/extract at that Step:
{
"steps": {
"pdf": {
"robot": "/document/convert",
"use": ":original",
"format": "pdf"
},
"text": {
"robot": "/document/extract",
"use": "pdf",
"extract": ["text"],
"text_format": "json"
}
}
}
Limits and pricing
page_range(for example"1,3-5") limits the work to the pages you need and can select up to 1,000 pages. JSON and per-page text output can cover at most 1,000 selected pages per job.page_rangeandpassword(for encrypted PDFs) apply to native text and embedded image extraction. Per-page text output requirestext_method: "native". OCR processes the whole document and cannot be combined with these options or open encrypted PDFs."auto"decides per document, not per page. If a PDF mixes typed and scanned pages, the native text wins and the scanned pages stay empty; use"ocr"for those files.- For OCR text, use
text_format: "txt". ThetextandpagesJSON shape shown above is for native extraction. - Vector graphics, charts, and page backgrounds are often not embedded images, so they are not
extracted. Use
/document/thumbsfor those.
Native extraction counts input and output bytes toward your plan, with a 1 MiB minimum per document.
The text Step in our nine-page example counted as that 1 MiB minimum. When OCR runs, per-page
charges are added; see the /document/ocr pricing.
Get started
/document/extract is available to every account. Read the
parameter reference, add the Step to a Template you already
use, or create a free account and run the example above. If your documents need
something the Robot does not do yet, let us know.
