Recognize text in images (OCR) in Rust
Build a Rust command-line program that takes an image filename and prints the text Tesseract recognizes. Start with a generated image containing a known invoice number, then try your own image. The same program includes an optional grayscale path and returns a nonzero exit status when it cannot recognize text.
This walkthrough uses local files and English language data. A correct result on the sample checks that the Rust bindings, native libraries, and model work together; it does not establish accuracy on receipts, handwriting, or complex page layouts.
Prerequisites
- Linux and Bash, with Rust and Cargo installed.
- Native Tesseract and Leptonica, including their development headers and libraries.
pkg-config, a C/C++ compiler, and libclang for the native bindings.- English language data,
eng.traineddata. - ImageMagick 7 and the Liberation Sans font for generating the sample image.
The example was run with Rust/Cargo 1.98.1, Tesseract 5.5.3, Leptonica 1.87.0, and ImageMagick
7.1.2-31. Use that toolchain for this walkthrough. The image crate’s own declared Rust minimum
is 1.88.0, but that is not a tested minimum
for this complete project.
Installing Tesseract
The Rust crate does not install Tesseract itself. Follow the native installation guide for your distribution and install the development libraries as well as the engine and English model.
On Ubuntu/Debian, the relevant package names include libtesseract-dev, libleptonica-dev, and
tesseract-ocr-eng. Distribution packages may provide versions different from the tested toolchain.
The runnable instructions here cover Linux, not a verified macOS or Windows setup. Check the native dependencies before creating the project:
rustc --version &&
cargo --version &&
tesseract --version &&
pkg-config --modversion tesseract lept &&
tesseract --list-langs &&
magick -version
The language list must include eng. If pkg-config cannot find tesseract or lept, resolve the
missing development package or its search path before compiling. ImageMagick’s font list
(magick -list font) should include Liberation-Sans.
Setting up the project
Choose a directory outside an existing Cargo workspace. Paste this block to create a new binary
project. It refuses an existing rust-ocr directory and stays in your original directory:
(cargo new --bin --edition 2021 --vcs none rust-ocr)
Inside rust-ocr, replace the complete Cargo.toml with:
[package]
name = "rust-ocr"
version = "0.1.0"
edition = "2021"
[dependencies]
anyhow = "=1.0.104"
image = { version = "=0.25.10", default-features = false, features = ["jpeg", "png"] }
tempfile = "=3.27.0"
tesseract = "=0.15.2"
The first build creates Cargo.lock. Keep it with your application and use --locked on subsequent
builds so Cargo refuses dependency resolution changes. The native libraries and language data are
installed separately; the lockfile does not pin them.
Basic OCR implementation
Replace rust-ocr/src/main.rs with this complete program:
use std::path::PathBuf;
use anyhow::{bail, ensure, Context, Result};
use tesseract::{PageSegMode, Tesseract};
const USAGE: &str = "Usage: rust-ocr <image> [--grayscale]";
fn main() -> Result<()> {
let mut arguments = std::env::args_os().skip(1);
let input = PathBuf::from(arguments.next().context(USAGE)?);
let grayscale = match arguments.next() {
None => false,
Some(flag) if flag == "--grayscale" => true,
_ => bail!(USAGE),
};
ensure!(arguments.next().is_none(), USAGE);
let input = input.canonicalize().context("Could not open input image")?;
let temporary_directory = if grayscale {
Some(tempfile::tempdir().context("Could not create temporary directory")?)
} else {
None
};
let image_path = match &temporary_directory {
Some(directory) => {
let decoded = image::open(&input).context("Could not decode input image")?;
ensure!(
matches!(decoded.color(), image::ColorType::L8 | image::ColorType::Rgb8),
"Grayscale mode requires opaque 8-bit PNG or JPEG"
);
let processed = directory.path().join("processed.png");
decoded.grayscale().save(&processed)
.context("Could not write grayscale image")?;
processed.canonicalize().context("Could not resolve temporary image")?
}
None => input,
};
let filename = image_path.to_str().context("Image path must be valid UTF-8")?;
let mut ocr = Tesseract::new(None, Some("eng"))
.context("Could not initialize English OCR; check eng.traineddata")?
.set_image(filename)
.context("Could not decode input image")?;
ocr.set_page_seg_mode(PageSegMode::PsmSingleBlock);
let text = ocr.get_text().context("Recognition failed")?;
ensure!(!text.trim().is_empty(), "No text recognized");
print!("{text}");
Ok(())
}
None lets Tesseract locate its installed model directory; Some("eng") explicitly selects
English. The binding API loads the
image with set_image and returns recognized UTF-8 text through get_text. Errors propagate out of
main instead of being printed and treated as success.
Run it on known text
From the directory containing rust-ocr, paste this block. Its subshell keeps your working
directory and shell options unchanged, including when a step fails:
(
cd rust-ocr &&
cargo build &&
magick -size 900x180 xc:white -font Liberation-Sans -pointsize 64 \
-fill black -gravity center -annotate +0+0 'INVOICE 12345' PNG24:input.png &&
cargo run --quiet --locked -- input.png
)
The ImageMagick annotation options create a readable, opaque 8-bit PNG. In the tested environment, stdout contains:
INVOICE 12345
This command replaces rust-ocr/input.png on a successful rerun. Reserve that filename for the
sample, not for an image you want to keep. If the build fails, the && chain does not generate the
image or run an older executable. Tesseract may also print diagnostics on stderr.
Handling different image formats
To exercise the grayscale path on the same sample:
(cd rust-ocr && cargo run --quiet --locked -- input.png --grayscale)
It produces the same recognized text. This path decodes PNG or JPEG through image, converts the
pixels to grayscale, and writes a PNG in its own temporary directory without modifying the input.
It deliberately rejects alpha-bearing and 16-bit images rather than silently discarding transparency
or changing sample ranges. Use upright, opaque 8-bit images for this walkthrough; other formats and
automatic orientation correction are outside its scope.
The TempDir stays alive until recognition finishes. Its destructor attempts cleanup on both
success and ordinary error returns. As the
tempfile source documents, forced
termination can leave files behind, and destructor cleanup errors are not reported. This is not a
secure-erasure guarantee.
Advanced OCR configuration
The program selects PsmSingleBlock, Tesseract’s mode 6, because the sample is one text block.
It does not request orientation/script detection, so this path needs eng but not osd data.
Choose segmentation based on the actual layout rather than assuming one mode fits all images;
Tesseract’s quality guide
explains the alternatives.
Grayscale conversion is an option to compare, not a promised accuracy improvement. Tesseract already does preprocessing internally. Blurred letters, skew, and low contrast still need better input or task-specific processing. A character whitelist would also exclude legitimate punctuation, so the example does not apply one.
Best practices for OCR in Rust
Treat a successful exit as “some text was recognized,” not “the text is correct.” Compare important values against the image before using them. Blank input is an error in this CLI, although an empty result can be valid behavior for an OCR engine. Try it explicitly:
(
cd rust-ocr &&
magick -size 900x180 xc:white PNG24:blank.png &&
cargo run --quiet --locked -- blank.png
)
This block replaces the sample blank.png and exits nonzero with
No text recognized. Other failures have different remedies:
- A missing filename or extra argument reports the usage message. Put
--grayscaleafter the image. - A nonexistent image reports Could not open input image.
- A decoder failure reports Could not decode input image.
- Missing English data reports Could not initialize English OCR; check eng.traineddata.
Decoding is not a file-integrity check. A damaged JPEG can still decode, particularly in the grayscale path, and yield incomplete text instead of a decoder error. Inspect the result against the original; this CLI does not certify that an image is intact.
If the English model is installed in a custom location, TESSDATA_PREFIX must name the existing
directory containing eng.traineddata, not the file itself. This is the lookup used by
Tesseract 5.5.3.
Check the directory listed by tesseract --list-langs; a stale environment override can point the
Rust program at a different model directory.
Handling multiple languages
This CLI deliberately selects only English. Installing additional models does not change that selection. Tesseract supports combined language identifiers, but a multilingual application also needs those models and representative images to evaluate its results. Consult the language configuration documentation when extending it; the English fixture does not verify multilingual recognition.
Conclusion
For your own image, replace input.png in the run command with its filename, quoting paths that
contain spaces. Keep the original image so you can compare the recognized text with it. If your next
step is a hosted document workflow rather than a local CLI, see
Transloadit’s document processing service.
