'unoconv' for document conversion ease
Converting documents between different formats is a common task for developers and professionals. Historically, unoconv served as an open-source tool that leveraged LibreOffice to convert documents such as DOCX, ODT, and XLSX to PDF and other formats.
Unoconv is deprecated and its repository is archived. For all new implementations, please use unoserver, the modern successor. This post is maintained for historical reference and for those supporting existing unoconv workflows.
Introduction to 'unoconv' and its open-source capabilities
'Unoconv' (Universal Office Converter) is a command-line utility that leverages LibreOffice for document conversion. It supports a wide range of document formats, making it useful for developers maintaining legacy systems or supporting existing workflows. Although unoconv is deprecated, it serves as an instructive example of automating document conversion tasks.
Setting up and installing 'unoconv' on your system
Prerequisites
Before using 'unoconv', ensure that LibreOffice is installed on your system since unoconv relies on it for performing conversions. Be aware that unoconv may have compatibility issues with modern Python versions; consider switching to unoserver for new projects.
Installation on Linux (Ubuntu/Debian)
For an existing legacy deployment, check whether your distribution still packages unoconv:
apt-cache policy libreoffice unoconv
Package availability does not guarantee compatibility with your Python and LibreOffice versions. Keep legacy environments pinned and tested. For a new deployment, follow the maintained unoserver installation instructions.
Installation on macOS
The unoconv Homebrew formula is disabled because its repository was archived. Do not use
brew install unoconv for a new setup. Install LibreOffice and follow the unoserver instructions
for selecting a Python interpreter that can import LibreOffice's uno module:
brew install --cask libreoffice
Installation on Windows
Native Windows support for unoconv is limited. We recommend using Windows Subsystem for Linux (WSL2) and following the Linux installation instructions to ensure compatibility.
Common use cases and scenarios
Developers can employ unoconv for various tasks, including:
- Batch Conversions: Convert multiple DOCX files to PDF in a single operation.
- Automating Report Generation: Automatically generate PDF reports from document templates.
- Legacy System Maintenance: Support and maintain workflows that still depend on unoconv.
- Server-Side Processing: Integrate document conversion into back-end applications for automated processing.
Step-by-step guide: converting documents from docx to PDF
Basic conversion using unoconv is straightforward. To convert a single DOCX file to PDF, execute:
unoconv -f pdf document.docx
For a small set in a directory that has no colliding PDF outputs:
unoconv -f pdf ./*.docx
Automating document conversion tasks with scripts
Automating document conversion can streamline your workflow. The following Python script
demonstrates how to convert DOCX files on an already working unoconv installation. Save it as
convert_documents.py and run python3 convert_documents.py docs new-pdfs. It requires a new output
directory, skips directories and symlinks, and publishes each result only after the command succeeds.
If any conversion fails, the process exits unsuccessfully after attempting the other files.
import os
from pathlib import Path
import subprocess
import sys
from tempfile import TemporaryDirectory
def convert_documents(input_dir, output_dir):
source = Path(input_dir).resolve(strict=True)
files = sorted(path for path in source.iterdir()
if not path.is_symlink() and path.is_file()
and path.suffix.lower() == '.docx')
if not files:
print('No DOCX files found', file=sys.stderr)
return 1
output = Path(output_dir).absolute()
output.mkdir(mode=0o700)
failures = 0
for path in files:
try:
with TemporaryDirectory(prefix='.conversion-', dir=output) as temporary:
candidate = Path(temporary) / 'result.pdf'
subprocess.run(
['unoconv', '-f', 'pdf', '-o', str(candidate), str(path)],
check=True, timeout=120,
)
if not candidate.is_file() or candidate.stat().st_size == 0:
raise ValueError('No nonempty PDF was produced')
# Hard-link publication fails instead of replacing a destination that appeared.
os.link(candidate, output / (path.stem + '.pdf'))
print(f'Converted {path.name}')
except (OSError, ValueError, subprocess.SubprocessError):
failures += 1
print(f'Conversion failed: {path.name}', file=sys.stderr)
return 1 if failures else 0
if __name__ == '__main__':
if len(sys.argv) != 3:
sys.exit('Usage: python3 convert_documents.py <input-directory> <new-output-directory>')
try:
sys.exit(convert_documents(sys.argv[1], sys.argv[2]))
except (OSError, ValueError):
sys.exit('Cannot prepare conversion directories; use an existing input and a new output directory')
The timeout bounds the client process, not necessarily work already running in LibreOffice. Use a dedicated conversion worker for untrusted files, with operating-system resource limits and no sensitive filesystem access. Check rendering, fonts and pagination on representative documents; a successful command is not a guarantee of visual fidelity.
Troubleshooting common issues
Python version conflicts
Modern systems might present Python version conflicts with unoconv. To resolve these issues, you can:
- Use a Python virtual environment with a compatible version.
- Install a specific Python version that works well with unoconv.
- Switch to using unoserver, which offers better Python 3 compatibility.
LibreOffice version compatibility
Ensure your LibreOffice installation is compatible with unoconv:
libreoffice --version
unoconv --version
Container deployment considerations
For containerized environments, consider the following:
- Build from a supported distribution and install its LibreOffice packages.
- Ensure necessary fonts are installed.
- Configure appropriate file permissions.
These steps help prevent errors during document conversion in container deployments.
Modern alternatives and comparison
Unoserver (recommended)
Unoserver is the successor to unoconv. It provides a long-running LibreOffice worker and a separate
unoconvert client. Its Python interpreter must be able to import uno; an arbitrary virtual
environment is not sufficient. For a complete integration, see
document conversion from Rust.
LibreOffice CLI
LibreOffice itself provides a native command-line conversion option that bypasses the need for unoconv. To convert a DOCX file to PDF using LibreOffice CLI, run:
soffice --headless --convert-to pdf document.docx
Cloud-based solutions
For scalable production environments, consider cloud-based document conversion services or containerized solutions that provide enhanced reliability and effortless maintenance.
Conclusion
Although unoconv has served developers well over the years, its deprecation means that new projects should consider modern alternatives like unoserver or cloud-based conversion services. For those maintaining legacy systems, be mindful of compatibility and potential Python version issues.
If you are seeking a robust, scalable solution for document conversion, consider exploring Transloadit's Document Conversion service.
