Java PDF processing with Ghost4J & Ghostscript
Ghostscript renders PDF and PostScript documents. This tutorial originally used Ghost4J, a Java wrapper around its native library. The updated examples run Ghostscript in separate processes and use Apache PDFBox for font inspection, avoiding Ghost4J’s obsolete dependency stack.
Introduction to Ghostscript
Ghostscript can render PDF pages to images and convert document formats. The Java code below launches its command-line interface directly with an argument list. Filenames never become shell commands, and each rendering job gets its own process and output directory.
What is Ghost4J?
Ghost4J wraps Ghostscript through native bindings. Its
published 1.0.1 POM
pins Log4j 1.2.17, JNA 4.1.0, and iText 2.1.7. Adding a Java executor around
a native singleton does not make it safe to run concurrently. This guide therefore replaces the
Ghost4J dependency rather than recommending that stack for new applications.
Separate processes isolate native interpreter state. They do not create a security sandbox: process
untrusted documents in workers with restricted filesystem access, no network access, and enforced
memory, CPU, and output-size limits. Keep Ghostscript patched; -dSAFER is an additional interpreter
restriction, not a substitute for those controls.
Set up Ghostscript and PDFBox in your Java project
Install a maintained JDK 21,
Ghostscript, and curl separately. The commands below
were tested on Linux with OpenJDK 21.0.12.1, Ghostscript 10.07.1, and PDFBox 3.0.8.
Later security updates within JDK 21 are appropriate; use a supported, patched Ghostscript build.
JDK 21 is the example’s chosen runtime, not PDFBox’s minimum requirement.
Work in one directory containing your local input.pdf, first.pdf, and second.pdf files. Save
all three Java programs below there. Check that java, javac, and gs run before downloading the
standalone PDFBox application JAR, which includes the font example’s dependencies:
java -version &&
javac -version &&
gs --version &&
curl --fail --location --output pdfbox-app-3.0.8.jar \
https://repo.maven.apache.org/maven2/org/apache/pdfbox/pdfbox-app/3.0.8/pdfbox-app-3.0.8.jar
For Maven applications, the corresponding library dependency is org.apache.pdfbox:pdfbox:3.0.8.
Do not add Ghost4J to run these examples. See the PDFBox downloads
for release checksums and the Ghostscript usage documentation
for rendering options. If a check or download fails, fix that prerequisite and rerun the setup block
before compiling. The commands use && so a failed setup or build stops the dependent command.
Convert PDF to image
Save this as PDFToImage.java. Pass the trusted Ghostscript executable path, input PDF, and a new
output directory. A fixed relative output pattern prevents percent characters in directory names
from becoming Ghostscript formatting directives.
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.TimeUnit;
public final class PDFToImage {
public static void render(Path executable, Path input, Path output) throws Exception {
Path source = input.toRealPath();
if (!Files.isRegularFile(source)) throw new IOException("Input must be a regular file");
Path binary = executable.toRealPath();
Files.createDirectory(output);
Path destination = output.toRealPath();
Process process = new ProcessBuilder(
binary.toString(), "-dSAFER", "-dBATCH", "-dNOPAUSE", "-dNOPROMPT",
"-sDEVICE=png16m", "-r144", "-sOutputFile=page-%03d.png", "-f", source.toString()
).directory(destination.toFile()).redirectErrorStream(true)
.redirectOutput(destination.resolve("ghostscript.log").toFile()).start();
try {
if (!process.waitFor(120, TimeUnit.SECONDS)) {
throw new IOException("Rendering timed out for " + source);
}
if (process.exitValue() != 0) {
throw new IOException("Rendering failed for " + source + "; inspect "
+ destination.resolve("ghostscript.log"));
}
try (var files = Files.list(destination)) {
if (files.noneMatch(path -> path.getFileName().toString().endsWith(".png"))) {
throw new IOException("No pages rendered for " + source);
}
}
} finally {
if (process.isAlive()) {
process.destroyForcibly();
process.waitFor();
}
}
}
public static void main(String[] args) throws Exception {
if (args.length != 3) throw new IllegalArgumentException("gs-path input.pdf output-directory");
render(Path.of(args[0]), Path.of(args[1]), Path.of(args[2]));
}
}
Compile and run it, substituting your installed executable path:
javac PDFToImage.java &&
java PDFToImage /usr/bin/gs input.pdf rendered
Open rendered/page-001.png; subsequent pages are numbered page-002.png, and so on, at 144 DPI.
A page measuring 144 × 72 PDF points produces a 288 × 144 pixel image. The output directory
must be new: rerunning with rendered already present fails and preserves its files. Choose another
directory to retry, retaining the failed job’s log and any partial images for diagnosis.
Rendering intentionally flattens the document; it does not preserve searchable text, forms, signatures, or accessibility structure. On failure, keep the output directory private for diagnosis and do not publish its partial images.
Concurrent PDF processing
Save this as ConcurrentPDFProcessing.java. Each worker launches its own Ghostscript process.
The executor waits for submitted work to finish, and a task failure reaches the command-line caller.
At most two rendering processes run together; this bounds active jobs, not the queued input list.
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.concurrent.Callable;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
public final class ConcurrentPDFProcessing {
public static void main(String[] args) throws Exception {
if (args.length < 3) throw new IllegalArgumentException("gs-path new-output-directory PDFs...");
Path output = Files.createDirectory(Path.of(args[1]));
try (ExecutorService executor = Executors.newFixedThreadPool(2)) {
var results = new ArrayList<Future<Void>>();
for (int i = 2; i < args.length; i++) {
Path input = Path.of(args[i]);
Path destination = output.resolve("document-" + (i - 2));
results.add(executor.submit((Callable<Void>) () -> {
PDFToImage.render(Path.of(args[0]), input, destination);
return null;
}));
}
for (Future<Void> result : results) result.get();
}
}
}
javac PDFToImage.java ConcurrentPDFProcessing.java &&
java ConcurrentPDFProcessing /usr/bin/gs batch-output first.pdf second.pdf
batch-output/document-0 contains the pages of first.pdf; document-1 contains those of
second.pdf. The batch directory must also be new. If any job fails, the command exits nonzero
after the remaining submitted jobs finish. Successful jobs can leave complete images beside a
failed job’s partial output; inspect the named input and log before using the batch.
Analyze fonts in PDF documents
PDFBox 3 loads files through Loader.loadPDF(). Resource font names are COSName keys; resolve
each key with PDResources.getFont(). Running a text stripper is unnecessary for enumerating these
resources. See the PDFBox migration guide.
Save this as FontAnalysis.java. It visits every page and nested Form XObject, prevents resource
cycles, and reports unique declared font names in sorted order, one per line. An unnamed font,
such as a Type 3 font without a name, is labeled by its resource key. This is a resource inventory,
not proof that every listed font draws visible text or a complete audit of fonts inside patterns
and Type 3 glyphs.
import java.io.IOException;
import java.nio.file.Path;
import java.util.Collections;
import java.util.IdentityHashMap;
import java.util.Set;
import java.util.TreeSet;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.cos.COSDictionary;
import org.apache.pdfbox.cos.COSName;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDResources;
import org.apache.pdfbox.pdmodel.graphics.form.PDFormXObject;
public final class FontAnalysis {
private static void collect(PDResources resources, Set<COSDictionary> visited,
Set<String> fonts) throws IOException {
if (resources == null || !visited.add(resources.getCOSObject())) return;
for (COSName name : resources.getFontNames()) {
var font = resources.getFont(name);
if (font == null) continue;
String fontName = font.getName();
fonts.add(fontName == null ? "Unnamed font (resource /" + name.getName() + ")" : fontName);
}
for (COSName name : resources.getXObjectNames()) {
if (resources.getXObject(name) instanceof PDFormXObject form) {
collect(form.getResources(), visited, fonts);
}
}
}
public static void main(String[] args) throws Exception {
if (args.length != 1) throw new IllegalArgumentException("input.pdf");
Set<COSDictionary> visited = Collections.newSetFromMap(new IdentityHashMap<>());
Set<String> fonts = new TreeSet<>();
try (PDDocument document = Loader.loadPDF(Path.of(args[0]).toFile())) {
for (var page : document.getPages()) collect(page.getResources(), visited, fonts);
}
for (String font : fonts) System.out.println(font);
}
}
Run the font example from the same directory:
javac -cp pdfbox-app-3.0.8.jar FontAnalysis.java &&
java -cp '.:pdfbox-app-3.0.8.jar' FontAnalysis input.pdf
Performance considerations and best practices
Resolution controls output dimensions and memory requirements. Choose a small worker count and measure representative files. Verify output page count, dimensions, and visible content; a zero exit status alone cannot detect a blank or clipped rendering. Test malformed PDFs as well as valid ones, and avoid logging document content or sharing private diagnostic logs.
Conclusion
If you would prefer a managed conversion workflow, see Transloadit’s document processing service.
