Building a document OCR tool using GCP OCR and Node.js
Extract text from a local PNG, JPEG, or PDF with one Node.js command-line program. Images go directly to Google Cloud Vision; PDFs are uploaded to a private Cloud Storage bucket, processed asynchronously, and returned as JSON with one text entry per page. The PDF command checks page completeness and job identity before printing its JSON result.
Prerequisites
- Node.js 26.8.2 or a newer Node.js 26 release, with native TypeScript support. Install the current security release from Node.js.
- Corepack installed separately, providing Yarn 4.12.0 through the project’s pinned package manager.
- The Google Cloud CLI.
- A PNG or JPEG you are allowed to send to Google Cloud. Put it at
sample.pngbesideocr-project, or adjust the image command’s path. - A Google Cloud project with billing, the Vision API, and Cloud Storage enabled. Your account needs permission to use the project’s quota, create and delete the example bucket, and create, list, read, and delete its objects. Ask your project administrator for access if needed.
The commands use Bash on macOS. The example was tested with Node.js 26.8.2, Yarn 4.12.0,
@google-cloud/vision 6.1.1, and @google-cloud/storage 8.2.0. This is a sequential local tutorial:
use synthetic documents first, and do not share a job prefix with another process.
Setting up Google Cloud Vision API
Select your project in the Google Cloud Console, confirm that
billing is enabled, and enable the Cloud Vision API. Google’s
Vision setup guide covers the project setup.
Replace PROJECT_ID consistently in the commands below.
Authentication setup
For local development, use Application Default Credentials (ADC) with an authorized user account. If both the CLI and ADC already use your intended account and quota project, skip these login commands:
gcloud auth login &&
gcloud auth application-default login &&
gcloud auth application-default set-quota-project PROJECT_ID
ADC is separate from the CLI’s normal login. The quota project requires serviceusage.services.use;
see Google’s Vision authentication guide.
An existing GOOGLE_APPLICATION_CREDENTIALS setting overrides the local ADC file, so check that it
selects your intended credentials. Do not download a service-account key for this walkthrough.
For a later deployment, use an attached identity or Workload Identity Federation with ADC.
Installing the Google Cloud Vision client library
Paste this block from the directory where you want the example. It checks the tools before writing,
refuses an existing ocr-project, and leaves your shell in its original directory:
(
set -eu
node --version
corepack --version
gcloud --version
test ! -e ocr-project && test ! -L ocr-project
mkdir ocr-project
cd ocr-project
cat > package.json <<'JSON'
{
"name": "node-document-ocr",
"private": true,
"type": "module",
"packageManager": "yarn@4.12.0",
"dependencies": {
"@google-cloud/storage": "8.2.0",
"@google-cloud/vision": "6.1.1",
"pdf-lib": "1.17.1",
"zod": "4.3.6"
}
}
JSON
touch yarn.lock
printf 'nodeLinker: node-modules\n' > .yarnrc.yml
YARN_IGNORE_PATH=1 corepack yarn install
)
The child lockfile and package-manager pin keep installation separate from an enclosing Yarn
project; the local configuration selects node_modules. If installation fails after project
creation, preserve the directory and retry only the installation:
(cd ocr-project && YARN_IGNORE_PATH=1 corepack yarn install)
If project creation failed, choose another parent directory before repeating setup. Save the two
TypeScript files below inside ocr-project; later commands still run from the parent directory.
Writing the Node.js code
Save this complete program as ocr-project/ocr.ts. It accepts an image path or a PDF path plus a
bucket name and job ID. PDF upload is part of the command. A preexisting job prefix is refused, and
the input upload also uses a generation precondition to refuse replacement.
import type { Bucket } from '@google-cloud/storage'
import { readFile } from 'node:fs/promises'
import { setTimeout as delay } from 'node:timers/promises'
import { Storage } from '@google-cloud/storage'
import { ImageAnnotatorClient, protos } from '@google-cloud/vision'
import { PDFDocument } from 'pdf-lib'
import { z } from 'zod'
const rpcOptions = { timeout: 30_000, retry: null }
const statusSchema = z.object({ code: z.number().optional() })
const pageSchema = z.object({
error: statusSchema.optional(),
context: z.object({ uri: z.string(), pageNumber: z.number().int() }),
fullTextAnnotation: z.object({ text: z.string().default('') }).optional(),
})
const fileSchema = z.object({
error: statusSchema.optional(),
inputConfig: z.object({ gcsSource: z.object({ uri: z.string() }) }),
responses: z.array(pageSchema),
})
const operationResultSchema = z.object({
responses: z.array(z.object({
outputConfig: z.object({ gcsDestination: z.object({ uri: z.string() }) }),
})).length(1),
})
class OcrError extends Error {}
function jobPrefix(job: string): string {
if (!/^[a-z0-9][a-z0-9-]{0,63}$/u.test(job)) {
throw new OcrError('Use a job ID containing lowercase letters, numbers, and hyphens')
}
return `node-ocr/${job}/`
}
async function removeJob(bucket: Bucket, prefix: string): Promise<void> {
const [files] = await bucket.getFiles({ prefix })
for (const file of files) await file.delete()
}
async function waitForOperation(
client: ImageAnnotatorClient,
name: string,
): Promise<protos.google.longrunning.IOperation> {
const deadline = Date.now() + 180_000
while (Date.now() < deadline) {
const [status] = await client.operationsClient.getOperation({ name }, rpcOptions)
if (status.done) return status
await delay(1000)
}
throw new OcrError('Wait expired; the cloud operation may still be running')
}
async function extractTextFromImage(
client: ImageAnnotatorClient,
bytes: Buffer,
): Promise<{ number: number; text: string }[]> {
const png = bytes.subarray(0, 8).equals(Buffer.from([137, 80, 78, 71, 13, 10, 26, 10]))
const jpeg = bytes[0] === 255 && bytes[1] === 216 && bytes[2] === 255
if (!png && !jpeg) throw new OcrError('Image input must be a PNG or JPEG')
const [result] = await client.documentTextDetection({ image: { content: bytes } }, rpcOptions)
if (result.error?.code) {
throw Object.assign(new OcrError('Vision rejected the image', { cause: result.error }), {
code: result.error.code,
})
}
return [{ number: 1, text: (result.fullTextAnnotation?.text ?? '').trimEnd() }]
}
async function extractTextFromPDF(
client: ImageAnnotatorClient,
bucket: Bucket,
bytes: Buffer,
job: string,
project: string,
): Promise<{ number: number; text: string }[]> {
const pdf = await PDFDocument.load(bytes)
const pageCount = pdf.getPageCount()
if (pageCount < 1 || pageCount > 20) throw new OcrError('This example accepts 1–20 PDF pages')
const prefix = jobPrefix(job)
const outputPrefix = `${prefix}output/`
const source = `gs://${bucket.name}/${prefix}input.pdf`
const destination = `gs://${bucket.name}/${outputPrefix}`
const [existing] = await bucket.getFiles({ prefix })
if (existing.length > 0) throw new OcrError('Job prefix already exists; choose a new job ID')
let uploaded = false
let safeToClean = true
try {
await bucket.file(`${prefix}input.pdf`).save(bytes, {
resumable: false,
contentType: 'application/pdf',
preconditionOpts: { ifGenerationMatch: 0 },
})
uploaded = true
// A lost submission response can leave an operation running without a known name.
safeToClean = false
const [operation] = await client.asyncBatchAnnotateFiles({
parent: `projects/${project}/locations/eu`,
requests: [{
inputConfig: { mimeType: 'application/pdf', gcsSource: { uri: source } },
features: [{ type: 'DOCUMENT_TEXT_DETECTION' }],
outputConfig: { batchSize: 1, gcsDestination: { uri: destination } },
}],
}, rpcOptions)
if (!operation.name) throw new OcrError('Vision did not return an operation name')
console.error(`Operation: ${operation.name}`)
await bucket.file(`${prefix}operation.json`).save(`${JSON.stringify({ name: operation.name })}\n`, {
resumable: false,
contentType: 'application/json',
preconditionOpts: { ifGenerationMatch: 0 },
})
const status = await waitForOperation(client, operation.name)
safeToClean = true
if (status.error?.code) {
throw Object.assign(new OcrError('Vision operation failed', { cause: status.error }), {
code: status.error.code,
})
}
if (!(status.response?.value instanceof Uint8Array)) throw new OcrError('Operation has no result')
const decoded = protos.google.cloud.vision.v1.AsyncBatchAnnotateFilesResponse.decode(
status.response.value,
)
const result = operationResultSchema.parse(decoded)
if (result.responses[0].outputConfig.gcsDestination.uri !== destination) {
throw new OcrError('Operation returned a different output prefix')
}
const [files] = await bucket.getFiles({ prefix: outputPrefix })
const pages = new Map<number, string>()
for (const file of files) {
if (!file.name.startsWith(outputPrefix) || !file.name.endsWith('.json')) {
throw new OcrError('Unexpected output object')
}
const [contents] = await file.download()
const data = fileSchema.parse(JSON.parse(contents.toString('utf8')))
if (data.error?.code) throw new OcrError('Vision file failed')
if (data.inputConfig.gcsSource.uri !== source) throw new OcrError('Wrong input in result')
for (const page of data.responses) {
if (page.error?.code) throw new OcrError('Vision page failed')
const number = page.context.pageNumber
if (page.context.uri !== source || number < 1 || number > pageCount || pages.has(number)) {
throw new OcrError('Wrong source, duplicate page, or invalid page number')
}
pages.set(number, (page.fullTextAnnotation?.text ?? '').trimEnd())
}
}
if (pages.size !== pageCount) throw new OcrError('Incomplete PDF result')
return Array.from({ length: pageCount }, (_, index) => ({
number: index + 1,
text: pages.get(index + 1) ?? '',
}))
} finally {
if (uploaded && safeToClean) await removeJob(bucket, prefix)
if (uploaded && !safeToClean) {
console.error(`Retained pending job: gs://${bucket.name}/${prefix}`)
}
}
}
async function main(): Promise<void> {
const [mode, input, bucketName, job, ...extra] = process.argv.slice(2)
if (!input || extra.length > 0 ||
(mode !== 'image' && mode !== 'pdf' && mode !== 'cleanup') ||
(mode === 'image' && (bucketName !== undefined || job !== undefined))) {
throw new OcrError('Usage: image FILE | pdf FILE BUCKET JOB | cleanup OPERATION BUCKET JOB')
}
const project = z.string().min(1).parse(process.env.GOOGLE_CLOUD_PROJECT)
const client = new ImageAnnotatorClient({ projectId: project, apiEndpoint: 'eu-vision.googleapis.com' })
const storage = new Storage({ projectId: project, retryOptions: { autoRetry: false }, timeout: 30_000 })
let pages: { number: number; text: string }[] | undefined
try {
if (mode === 'image') {
pages = await extractTextFromImage(client, await readFile(input))
} else {
if (!bucketName || !job) throw new OcrError('PDF and cleanup modes require BUCKET and JOB')
if (mode === 'cleanup') {
if (!input.startsWith(`projects/${project}/locations/eu/operations/`)) {
throw new OcrError('Use the operation name printed by this project’s PDF command')
}
const prefix = jobPrefix(job)
const bucket = storage.bucket(bucketName)
const [receipt] = await bucket.file(`${prefix}operation.json`).download()
const record = z.object({ name: z.string() }).parse(JSON.parse(receipt.toString('utf8')))
if (record.name !== input) throw new OcrError('Operation does not belong to this job prefix')
await waitForOperation(client, input)
await removeJob(bucket, prefix)
console.error('Job objects removed')
} else {
pages = await extractTextFromPDF(client, storage.bucket(bucketName), await readFile(input), job, project)
}
}
} finally {
await client.close()
}
if (pages !== undefined) console.log(JSON.stringify({ pages }, null, 2))
}
main().catch((error: unknown) => {
const code = statusSchema.safeParse(error)
const hint = code.success && code.data.code === 7 ? 'Permission denied; check project and bucket access'
: code.success && code.data.code === 8 ? 'Quota exceeded; check Vision quota'
: error instanceof OcrError ? error.message : 'Check input, ADC, bucket access, and result format'
console.error(`OCR failed: ${hint}`)
process.exitCode = 1
})
Run image mode with a local image that you are allowed to send to Google Cloud:
GOOGLE_CLOUD_PROJECT=PROJECT_ID node ocr-project/ocr.ts image ./sample.png
It prints one page containing the recognized text, without trailing whitespace. A blank image
returns an empty text string. A missing file, rejected image, or failed request exits with status
one and no JSON result. The image header check selects the format; Vision still has to decode the
image. Review OCR text against the original before using it as an authoritative transcription.
Processing PDF files
Save this fixture generator as ocr-project/make-pdf.ts. It creates two pages with different text
and refuses to overwrite an existing sample.pdf:
import { writeFile } from 'node:fs/promises'
import { PDFDocument, StandardFonts } from 'pdf-lib'
const pdf = await PDFDocument.create()
const font = await pdf.embedFont(StandardFonts.Helvetica)
for (const text of ['DEVTIPS FIRST PAGE 2026', 'DEVTIPS SECOND PAGE 2026']) {
const page = pdf.addPage([1000, 400])
page.drawText(text, { x: 50, y: 250, size: 32, font })
}
await writeFile(new URL('./sample.pdf', import.meta.url), await pdf.save(), { flag: 'wx' })
Generate it, then create a fresh private bucket. Replace YOUR_NEW_BUCKET with a globally unique
bucket name you own. The account used by gcloud must also be authorized for this project:
node ocr-project/make-pdf.ts &&
gcloud storage buckets create gs://YOUR_NEW_BUCKET --project=PROJECT_ID --location=EU \
--uniform-bucket-level-access --public-access-prevention &&
GOOGLE_CLOUD_PROJECT=PROJECT_ID node ocr-project/ocr.ts pdf \
ocr-project/sample.pdf YOUR_NEW_BUCKET trial-1
The CLI uploads node-ocr/trial-1/input.pdf and requests JSON outputs beneath
node-ocr/trial-1/output/. It waits for the operation, checks the returned destination and each
page’s source URI, rejects missing or duplicate pages, and orders the text by page number. Listing
objects alphabetically is insufficient for page ordering. Google’s
PDF/TIFF OCR guide explains the asynchronous handoff
and JSON response format.
For this fixture, the JSON result should be:
{
"pages": [
{ "number": 1, "text": "DEVTIPS FIRST PAGE 2026" },
{ "number": 2, "text": "DEVTIPS SECOND PAGE 2026" }
]
}
The successful command removes its uploaded input and OCR objects before printing the result. Terminal operation errors and invalid results also trigger that cleanup. A cleanup failure returns status one instead of publishing a successful result. Existing prefix contents are preserved when the command refuses a job; choose a new job ID rather than deleting somebody else’s objects.
Recover a pending job and remove the bucket
The CLI polls for up to three minutes; an in-flight status request can add up to 30 seconds. A timeout or lost connection does not cancel Vision. The command prints the operation name and retains the job objects when completion is unknown. Do not submit another OCR request to recover it. Copy the printed operation name into this command, using the same bucket and job ID:
GOOGLE_CLOUD_PROJECT=PROJECT_ID node ocr-project/ocr.ts cleanup \
'projects/PROJECT_ID/locations/eu/operations/OPERATION_ID' YOUR_NEW_BUCKET trial-1
Cleanup checks the saved operation record, waits for terminal status, and then removes that job’s objects, including after a failed operation. If it still times out, leave the objects and repeat cleanup later. If the submission response was lost before an operation name arrived, or saving the operation record failed, retain the prefix and have your project administrator verify the operation before deleting its objects. Only use cleanup on a prefix you created.
After your own example bucket is empty, remove it:
gcloud storage buckets delete gs://YOUR_NEW_BUCKET --project=PROJECT_ID --quiet
If fixture generation or bucket creation failed, preserve your files and retry the failed step alone. The generator refuses an existing PDF; reuse that file for OCR, or deliberately remove only that generated fixture before regenerating it. Do not run bucket deletion unless your creation step succeeded. Retain the JSON separately if you need it; this example does not maintain an archive.
API limitations and pricing
Vision accepts PDF/TIFF files up to 2,000 pages; this CLI deliberately accepts only PDFs with one to 20 pages. It does not implement TIFF. Each PDF page is billed as an image for the requested feature, so the two-page fixture uses two OCR units. Consult Cloud Vision pricing for current rates; Storage charges are separate. Request submission has no automatic retries, to avoid silently resubmitting a paid job after an uncertain response.
Regional configuration
This recipe uses the EU Vision endpoint, an EU request parent, and an EU Storage bucket. The endpoint controls Vision processing; it does not relocate an existing bucket. See Google’s multi-regional OCR documentation before adapting the location. The walkthrough was verified in the EU configuration.
Troubleshooting
- Authentication or quota errors: confirm the intended ADC account, quota project, and API
enablement. A working
gcloudlogin alone does not prove SDK authentication. - Bucket access errors: check input read, output create, list, read, and delete permissions. A denied download must not be treated as an empty OCR result.
- Incomplete or mismatched PDF output: keep the failure visible. Do not weaken the page count or source checks to publish partial text.
- An interrupted or timed-out job: recover its operation as above before removing resources or considering a fresh submission.
Conclusion
For an upload-processing workflow managed through Transloadit, see the Document OCR Robot. Its setup and result contract are separate from this direct Google Cloud CLI.
