Integrating OCR in the browser with Tesseract.js
Select a local image, click “Recognize text,” and copy its text without uploading the image. This walkthrough gives web developers a complete Tesseract.js example with loading feedback, a same-file retry, and one recognition job at a time. Start with a clear screenshot of printed English text; OCR output still needs proofreading.
What runs in the browser
Tesseract.js runs the Tesseract OCR engine through WebAssembly in a worker. The browser does the recognition, so speed and memory use depend on the reader’s device and image. There is no promise of an instant result. The project’s scope also rules out direct PDF input: render PDF pages to images separately before recognizing them.
This example pins Tesseract.js and its core to 7.0.0 and English language data to 1.0.0. It uses
text output, which is enabled by default. The worker API
initializes the language in the asynchronous createWorker('eng', 1, options) call; older
loadLanguage() and initialize() steps are unnecessary.
Browser compatibility and requirements
Use a current browser with Web Workers, nested workers, and WebAssembly. The complete example was tested in Chromium 145 and 152 on Linux. You also need Python 3 to serve the two files locally, plus a network connection to load the pinned scripts, WASM, and language data from jsDelivr. No Node.js build or package installation is needed.
The image stays in the browser in this example, but the asset downloads still contact a CDN. Language caching alone does not make the page work offline. An offline application must also serve or cache its HTML, scripts, worker, WASM, and language assets; this walkthrough does not install an offline cache. See the project’s asset hosting options.
Getting started with Tesseract.js
Installation
Create a new, empty directory named tesseract-browser. If that name already exists, choose
another directory instead of replacing its files. Save the next two blocks as index.html
and ocr-worker.js inside it. These are plain browser files, with no framework or backend.
Basic example: recognizing text from an image
Save this as index.html. The file picker and button stay disabled while a job is pending.
Selecting another image afterward clears the previous result; clicking the button again retries
the selected file without requiring a new file-selection event.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>Read text from a local image</title>
</head>
<body>
<h1>Read text from a local image</h1>
<form id="ocrForm">
<fieldset id="controls">
<legend>Recognize English text</legend>
<label for="imageInput">Image (JPEG, PNG, or WebP; up to 5 MiB)</label>
<input id="imageInput" type="file" accept="image/jpeg,image/png,image/webp" />
<button type="submit">Recognize text</button>
</fieldset>
</form>
<p id="status" role="status">Choose an image to begin.</p>
<label for="result">Recognized text</label>
<textarea id="result" rows="12" cols="60" readonly></textarea>
<script>
const form = document.getElementById('ocrForm')
const controls = document.getElementById('controls')
const imageInput = document.getElementById('imageInput')
const status = document.getElementById('status')
const result = document.getElementById('result')
let busy = false
async function recognizeImage(file) {
let task
let timer
let timedOut = false
try {
task = new Worker('./ocr-worker.js')
return await new Promise((resolve, reject) => {
const fail = () => reject(new Error('OCR task failed.'))
timer = setTimeout(() => {
timedOut = true
fail()
}, 90_000)
task.onerror = (event) => {
event.preventDefault()
fail()
}
task.onmessage = ({ data }) => {
if (data.type === 'result') resolve(data.text)
else if (data.type === 'error') fail()
else if (data.type === 'progress') {
status.textContent = data.status === 'recognizing text'
? 'Recognizing text… ' + Math.round(data.progress * 100) + '%'
: 'Loading OCR assets…'
}
}
task.postMessage(file)
})
} catch {
throw new Error(timedOut
? 'OCR timed out after 90 seconds. Try a smaller image or retry.'
: 'OCR failed. Check your connection and image, then try again.')
} finally {
clearTimeout(timer)
task?.terminate()
}
}
async function validateAndPerformOCR(file) {
if (!file || !['image/jpeg', 'image/png', 'image/webp'].includes(file.type)) {
throw new Error('Choose a JPEG, PNG, or WebP image.')
}
if (file.size === 0 || file.size > 5 * 1024 * 1024) {
throw new Error('Choose a nonempty image of 5 MiB or smaller.')
}
return recognizeImage(file)
}
imageInput.addEventListener('change', () => {
if (busy) return
result.value = ''
status.textContent = 'Selection changed. Click Recognize text.'
})
form.addEventListener('submit', async (event) => {
event.preventDefault()
if (busy || controls.disabled) return
const file = imageInput.files[0]
busy = true
controls.disabled = true
result.value = ''
status.textContent = 'Loading OCR assets…'
try {
const text = await validateAndPerformOCR(file)
result.value = text
status.textContent = text.trim()
? 'Finished: ' + file.name + '. You can copy the text below.'
: 'No text found. Try a clearer image of printed text.'
} catch (error) {
status.textContent = error instanceof Error
? error.message
: 'OCR failed. Try another image.'
} finally {
busy = false
controls.disabled = false
}
})
if (typeof Worker === 'undefined' || typeof WebAssembly === 'undefined') {
controls.disabled = true
status.textContent = 'Use a browser with Web Workers and WebAssembly.'
}
</script>
</body>
</html>
Save this as ocr-worker.js. The page owns this outer worker and can terminate it even if
Tesseract never finishes initializing. That matters because a failed language download can leave
createWorker() pending in version 7.0.0.
Its errorHandler reports failure directly to the page; the 90-second deadline also covers a
stalled download. The deadline is a demo policy, not an expected recognition time.
importScripts('https://cdn.jsdelivr.net/npm/tesseract.js@7.0.0/dist/tesseract.min.js')
async function performOCR(file) {
const worker = await Tesseract.createWorker('eng', 1, {
workerPath: 'https://cdn.jsdelivr.net/npm/tesseract.js@7.0.0/dist/worker.min.js',
corePath: 'https://cdn.jsdelivr.net/npm/tesseract.js-core@7.0.0',
langPath: 'https://cdn.jsdelivr.net/npm/@tesseract.js-data/eng@1.0.0/4.0.0_best_int',
logger: ({ status, progress }) => postMessage({ type: 'progress', status, progress }),
errorHandler: () => postMessage({ type: 'error' }),
})
try {
const { data: { text } } = await worker.recognize(file)
return text
} finally {
await worker.terminate()
}
}
self.onmessage = async ({ data: file }) => {
try {
const text = await performOCR(file)
postMessage({ type: 'result', text })
} catch {
postMessage({ type: 'error' })
}
}
From the parent of tesseract-browser, start a local server in a terminal:
(cd tesseract-browser && python3 -m http.server --bind 127.0.0.1 0)
Port 0 asks the OS for an available port. Open the http://127.0.0.1:PORT/ address printed
by the server. Use HTTP instead of double-clicking the HTML file, since worker loading depends on
the page’s origin. Stop the server with Ctrl+C when finished. Starting it again serves the same
files without overwriting them.
Choose a small screenshot containing “BROWSER OCR TEST”, then click “Recognize text.” You should see loading feedback, recognition progress, and the extracted words in “Recognized text.” A blank image should instead report “No text found.” The original file is never modified, and the page does not save images or recognized text to storage.
Error handling and validation
The five-MiB limit is this demo’s input policy, not Tesseract’s maximum. Compressed file size does
not bound decoded pixel memory, so begin with small images on mobile devices. The MIME allowlist
helps catch a wrong selection; a corrupt file labeled image/png still needs to fail during
recognition. Neither an image MIME type nor the picker’s
accept attribute
proves valid content.
If OCR fails, check the image and the browser’s Network panel for failed script, WASM, or language requests, then click “Recognize text” again. You can retry the same file. A blank result is not a worker error, and it does not prove the source image contains no text. Low contrast, tiny letters, and the wrong recognition language can also yield empty or inaccurate output.
Handling multiple languages
For mixed English and German text, add this function to ocr-worker.js and replace the handler’s
performOCR(file) call with performMultilingualOCR(file). The language array selects models;
it does not translate their output. This variation uses Tesseract’s default language URLs so each
language can load its own data, rather than the main example’s pinned English-only path.
async function performMultilingualOCR(file, languages = ['eng', 'deu']) {
const worker = await Tesseract.createWorker(languages, 1, {
workerPath: 'https://cdn.jsdelivr.net/npm/tesseract.js@7.0.0/dist/worker.min.js',
corePath: 'https://cdn.jsdelivr.net/npm/tesseract.js-core@7.0.0',
logger: ({ status, progress }) => postMessage({ type: 'progress', status, progress }),
errorHandler: () => postMessage({ type: 'error' }),
})
try {
const { data: { text } } = await worker.recognize(file)
return text
} finally {
await worker.terminate()
}
}
Performance optimization
Image preprocessing
First compare results on your actual images. Cropping excessive borders or correcting rotation can help; shrinking letters or increasing contrast can remove useful detail. The optional helper below caps width at 1,000 pixels as a memory tradeoff, not an accuracy recommendation.
Add it inside the index.html script and replace return recognizeImage(file) in
validateAndPerformOCR with return optimizedOCR(file). Keep validation before preprocessing.
async function preprocessImage(file) {
const url = URL.createObjectURL(file)
try {
const img = new Image()
await new Promise((resolve, reject) => {
img.onload = resolve
img.onerror = () => reject(new Error('Unable to decode image.'))
img.src = url
})
const canvas = document.createElement('canvas')
const maxWidth = 1000
const scale = img.width > maxWidth ? maxWidth / img.width : 1
canvas.width = Math.max(1, Math.round(img.width * scale))
canvas.height = Math.max(1, Math.round(img.height * scale))
const ctx = canvas.getContext('2d')
if (!ctx) throw new Error('Canvas processing is unavailable.')
ctx.filter = 'grayscale(100%) contrast(150%)'
ctx.drawImage(img, 0, 0, canvas.width, canvas.height)
return await new Promise((resolve, reject) => {
canvas.toBlob((blob) => {
if (blob) resolve(blob)
else reject(new Error('Unable to encode processed image.'))
}, 'image/png')
})
} finally {
URL.revokeObjectURL(url)
}
}
async function optimizedOCR(file) {
const processedImage = await preprocessImage(file)
return recognizeImage(processedImage)
}
Memory management
The single-image demo creates a fresh worker per attempt to keep its lifetime simple. For a batch,
reuse one initialized Tesseract worker and recognize the images sequentially. This preserves input
order and avoids loading a separate OCR engine for every image. The function rejects at the first
failed image and terminates its initialized worker in finally.
Add this to ocr-worker.js. For a minimal two-pass experiment with the current page, replace the
handler’s const text = await performOCR(file) with
const text = (await batchProcessImages([file, file])).join('\n'). The output contains two copies
in order; in an application, pass your ordered array of image files instead. The outer deadline
still applies to the entire job, so choose a suitable batch limit before expanding the UI.
async function batchProcessImages(files) {
const worker = await Tesseract.createWorker('eng', 1, {
workerPath: 'https://cdn.jsdelivr.net/npm/tesseract.js@7.0.0/dist/worker.min.js',
corePath: 'https://cdn.jsdelivr.net/npm/tesseract.js-core@7.0.0',
langPath: 'https://cdn.jsdelivr.net/npm/@tesseract.js-data/eng@1.0.0/4.0.0_best_int',
logger: ({ status, progress }) => postMessage({ type: 'progress', status, progress }),
errorHandler: () => postMessage({ type: 'error' }),
})
const results = []
try {
for (const file of files) {
const { data: { text } } = await worker.recognize(file)
results.push(text)
}
return results
} finally {
await worker.terminate()
}
}
Security considerations and best practices
Keep recognized text as text. This example assigns it to a read-only textarea’s value, so
recognized HTML cannot execute. If you move the result into another view, do not insert it through
innerHTML.
Local recognition describes where this code processes the image, not a blanket privacy guarantee
for any page that embeds it. Scripts on a page can access selected files. Review those dependencies
and your site’s other scripts before handling sensitive documents. If you host the OCR assets
yourself, serve WASM with application/wasm and keep the full matching core package so Tesseract
can select a build supported by the device.
