Automate document conversion in Rust with unoserver
Convert one document or a folder to PDFs by having Rust call unoserver’s unoconvert client. The
CLI below keeps full input filenames, refuses existing output directories, and returns a nonzero
exit if any selected file fails. LibreOffice stays running in a loopback worker while Rust attempts
files sequentially.
Install unoserver and Rust dependencies
Use Linux with LibreOffice’s Writer, Calc, and Impress components, its matching Python UNO bridge,
and Python’s venv support installed. On Debian/Ubuntu, these come from the libreoffice,
python3-uno, and python3-venv packages. Install a maintained Rust toolchain with Cargo using the
Rust installation instructions.
The tested versions were Rust/Cargo 1.98.1, Python 3.14.7, LibreOffice 26.8.0.3, and unoserver 3.7;
the compiler version is a test record, not an exact pin required by this example.
Create a new project without Cargo dependencies. The empty [workspace] table makes it its own
workspace root,
so setup does not add a member to an enclosing project’s manifest. The subshell leaves your
terminal’s directory unchanged, and mkdir refuses an existing project:
(
command -v cargo >/dev/null &&
command -v rustc >/dev/null &&
mkdir doc_converter_rust &&
cd doc_converter_rust &&
mkdir src &&
cat > Cargo.toml <<'TOML'
[package]
name = "doc_converter_rust"
version = "0.1.0"
edition = "2021"
[workspace]
TOML
)
From the directory where you created the project, run cd doc_converter_rust in each terminal.
Use that project directory for the remaining commands. Install the pinned client and server into
this project’s venv:
python3 -m venv --system-site-packages .venv &&
.venv/bin/python -c 'import uno' &&
.venv/bin/python -m pip install unoserver==3.7 &&
.venv/bin/unoconvert --version
The UNO import must pass before installation continues. If it fails, install the bridge for that
exact Python interpreter, then repeat this block; it reuses the partial venv without replacing your
project files. A default isolated pipx environment may not see the bridge. The
unoserver installation guide
explains this Python/LibreOffice pairing and lists macOS and Windows as unsupported; this walkthrough
covers Linux.
In one terminal, run the worker with both interfaces bound to loopback and separate ports:
.venv/bin/python - <<'PY' &&
import socket
for port in (2003, 2002):
with socket.socket() as listener:
listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
try:
listener.bind(('127.0.0.1', port))
except OSError as error:
raise SystemExit(f'Loopback port {port} is unavailable: {error}')
PY
.venv/bin/unoserver --interface 127.0.0.1 --port 2003 \
--uno-interface 127.0.0.1 --uno-port 2002 --conversion-timeout 120
Keep that terminal open. After startup, check the worker from the second terminal:
.venv/bin/unoping --host 127.0.0.1 --port 2003
Expect version information, including unoserver 3.7 and LibreOffice. The port check refuses existing listeners before starting LibreOffice. If it fails, choose a free pair and update the probe, server flags, and Rust client’s port below to match. Do not connect to an unidentified existing listener. The check releases its sockets before startup; if a listener appears in that gap and startup fails, stop this worker’s remaining LibreOffice child before retrying. Both interfaces lack authentication, and the conversion timeout terminates LibreOffice and exits the server if a conversion hangs, as described in the client/server documentation. Keep both ports private. Press Ctrl+C in the worker terminal to stop it when you finish.
Integrate Rust with unoserver
Save this complete CLI as src/main.rs. It accepts one file or a directory and requires a new
output directory. Each PDF retains the full input filename plus .pdf, so report.doc and
report.docx cannot collide. Filesystem paths remain native paths rather than being converted to
lossy strings for process arguments.
use std::error::Error;
use std::ffi::OsStr;
use std::fs::{self, DirBuilder, File};
use std::io::{self, Read};
use std::os::unix::fs::DirBuilderExt;
use std::path::{Path, PathBuf};
use std::process::{Command, ExitCode};
type Result<T> = std::result::Result<T, Box<dyn Error>>;
fn supported(path: &Path) -> bool {
let extension = path.extension().and_then(OsStr::to_str).unwrap_or("").to_ascii_lowercase();
matches!(extension.as_str(), "doc" | "docx" | "odt" | "rtf" | "txt" |
"ppt" | "pptx" | "odp" | "xls" | "xlsx" | "ods" | "csv")
}
fn convert_one(input: &Path, output: &Path) -> Result<()> {
let candidate = output.join(".conversion.pdf");
let conversion = (|| -> Result<()> {
let status = Command::new("unoconvert")
.args(["--host", "127.0.0.1", "--port", "2003", "--host-location", "local"])
.arg(input).arg(&candidate).status()?;
if !status.success() {
return Err(io::Error::other("unoconvert failed").into());
}
let mut header = [0; 5];
File::open(&candidate)?.read_exact(&mut header)?;
if &header != b"%PDF-" {
return Err(io::Error::other("conversion did not produce a PDF").into());
}
let mut filename = input.file_name().ok_or_else(|| io::Error::other("missing filename"))?.to_os_string();
filename.push(".pdf");
// Same-filesystem publication must fail rather than overwrite a destination.
fs::hard_link(&candidate, output.join(filename))?;
Ok(())
})();
if candidate.exists() {
fs::remove_file(&candidate)?;
}
conversion
}
fn run() -> Result<()> {
let args: Vec<_> = std::env::args_os().skip(1).collect();
if args.len() != 2 {
return Err(io::Error::other("Usage: doc_converter_rust <input-file-or-directory> <new-output-directory>").into());
}
let source = PathBuf::from(&args[0]);
let metadata = fs::symlink_metadata(&source)?;
let mut inputs = Vec::new();
if metadata.is_file() && supported(&source) {
inputs.push(source.canonicalize()?);
} else if metadata.is_dir() {
for entry in fs::read_dir(&source)? {
let entry = entry?;
if entry.file_type()?.is_file() && supported(&entry.path()) {
inputs.push(entry.path().canonicalize()?);
}
}
} else {
return Err(io::Error::other("use a supported regular file or directory, not a symlink").into());
}
inputs.sort();
if inputs.is_empty() {
return Err(io::Error::other("no supported documents found").into());
}
let output = PathBuf::from(&args[1]);
DirBuilder::new().mode(0o700).create(&output)?;
let output = output.canonicalize()?;
let mut failures = 0;
for input in inputs {
match convert_one(&input, &output) {
Ok(()) => println!("Converted {}", input.display()),
Err(_) => {
failures += 1;
eprintln!("Conversion failed for {}", input.display());
}
}
}
if failures > 0 {
return Err(io::Error::other(format!("{failures} conversion(s) failed")).into());
}
Ok(())
}
fn main() -> ExitCode {
match run() {
Ok(()) => ExitCode::SUCCESS,
Err(error) => {
eprintln!("Document conversion failed: {error}");
ExitCode::FAILURE
}
}
}
With a real DOCX file at the input path, run:
PATH="$PWD/.venv/bin:$PATH" cargo run -- /absolute/path/example.docx new-pdfs
Expect a Converted line naming the input and new-pdfs/example.docx.pdf. Open that PDF and check
its text, page count, and layout against the source. The %PDF- check only verifies the header;
it cannot prove that every page or embedded object survived conversion. LibreOffice may also
detect text renamed .docx, so a renamed file is not a reliable test of conversion failure.
Run the CLI and server as the same restricted account with access to the same filesystem, as
required by --host-location local. The output filesystem must support
hard links; creating the final link
fails if that filename already exists.
Batch-convert whole folders
The same executable handles a non-recursive folder batch without launching unbounded tasks:
PATH="$PWD/.venv/bin:$PATH" cargo run -- /absolute/path/documents new-batch-pdfs
It ignores subdirectories, symlinks, and unsupported extensions, attempts supported files in sorted order, preserves successful results, and exits unsuccessfully if any conversion fails. A file extension is only an initial filter, not proof that the file is valid or safe. Do not let other processes modify the input or output directories during a job.
For report.doc and report.docx, expect separate report.doc.pdf and report.docx.pdf files.
If another selected file fails, these successful PDFs remain, the failing filename appears in
stderr, and the batch exits with status 1. Retry into a new output directory after fixing the
input or worker; repeating the command with the old directory fails without replacing its PDFs.
Handle errors gracefully
- A missing input, empty selection, or existing output directory is a visible failure.
- A failed conversion does not publish its partial PDF. Inspect the nonzero batch status even if some earlier conversions succeeded.
- A server-side timeout exits the worker. Restart it before retrying, and keep retry counts bounded.
- The client is not a supervisor: add a whole-job deadline that terminates the CLI and its child process. Killing only the client does not necessarily cancel work already running in LibreOffice.
- Local diagnostics can include paths or document details. Keep them in restricted logs rather than returning raw converter errors through a web API.
The supplied CLI targets one worker. If an application queues uploads, keep the worker behind that application’s authentication and storage checks; exposing XML-RPC lets callers bypass them. For additional workers, configure separate LibreOffice profiles and non-overlapping XML-RPC/UNO port pairs, then route each job explicitly rather than starting a server for every conversion.
Supported formats at a glance
Actual support depends on the installed LibreOffice components and filters. Start with this conservative selection and verify representative documents:
| Category | Example inputs | Output in this CLI |
|---|---|---|
| Word processing | DOCX, DOC, ODT, RTF, TXT | |
| Spreadsheets | XLSX, XLS, ODS, CSV | |
| Presentations | PPTX, PPT, ODP |
Install the required fonts and inspect pagination, formulas, embedded objects, and rendering. A
successful process and a %PDF- file-header check do not establish visual fidelity or validate a
digital signature.
For managed document conversion, see Transloadit’s 🤖 /document/convert Robot.
