Read files in Java with memory mapping
Java’s FileChannel.map() lets you read file bytes through a MappedByteBuffer. To try it on a
concrete task, the program below computes a file’s SHA-256 checksum using either memory mapping or
buffered reads. Both modes process the same bytes, so you can check correctness before comparing
their performance on your files.
Understanding memory-mapped files
A mapping exposes a region of a file through virtual memory. Accessing its bytes can require the operating system to bring file pages into RAM; mapping a file does not mean that every byte is already resident. Java’s direct-buffer documentation describes mapped buffers as direct buffers whose contents live outside the ordinary Java heap.
This example hashes the bytes without decoding them. Text, UTF-8 characters, and binary data all
follow the same path. There is no encoding validation or newline normalization. If you instead
decode an entire mapping into a String, the decoded characters and resulting string still require
heap space proportional to the text size. Mapping does not remove that allocation.
Important limitations
Use an unchanging regular file and a 64-bit JDK 21. The commands below use Bash and were tested on Linux with OpenJDK 21.0.12.1; Windows and macOS file behavior has not been tested here.
The three-argument FileChannel.map()
accepts at most Integer.MAX_VALUE bytes per region: 2,147,483,647 bytes, or 2 GiB minus one byte.
The mapped mode below rejects larger files. Its buffered counterpart can read larger files with a
reusable 64 KiB buffer. Multiple mappings are another option, but this example uses one region to
keep its lifetime and size limit visible.
Do not modify or truncate the input during either read. A read-only mapping does not freeze the
file, and truncation can make mapped bytes inaccessible, as the
MappedByteBuffer contract
explains. Neither mode provides a snapshot of a file being written by another process.
Using Java NIO for memory-mapped files
Save this as FileChecksum.java in a working directory of your choice. It needs no dependencies
beyond the JDK.
import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.MappedByteBuffer;
import java.nio.channels.FileChannel;
import java.nio.file.Path;
import java.nio.file.StandardOpenOption;
import java.security.MessageDigest;
import java.security.NoSuchAlgorithmException;
import java.util.HexFormat;
public class FileChecksum {
private static byte[] checksum(Path path, boolean mapped)
throws IOException, NoSuchAlgorithmException {
MessageDigest digest = MessageDigest.getInstance("SHA-256");
try (FileChannel channel = FileChannel.open(path, StandardOpenOption.READ)) {
long size = channel.size();
if (size == 0) {
return digest.digest();
}
if (mapped) {
if (size > Integer.MAX_VALUE) {
throw new IllegalArgumentException(
"Mapped mode supports at most 2147483647 bytes; use buffered mode");
}
MappedByteBuffer buffer = channel.map(FileChannel.MapMode.READ_ONLY, 0, size);
digest.update(buffer);
} else {
ByteBuffer buffer = ByteBuffer.allocate(64 * 1024);
while (channel.read(buffer) != -1) {
buffer.flip();
digest.update(buffer);
buffer.clear();
}
}
}
return digest.digest();
}
public static void main(String[] args) throws IOException, NoSuchAlgorithmException {
if (args.length != 2 || !(args[0].equals("mapped") || args[0].equals("buffered"))) {
throw new IllegalArgumentException(
"Usage: java FileChecksum <mapped|buffered> <path>");
}
byte[] result = checksum(Path.of(args[1]), args[0].equals("mapped"));
System.out.println(HexFormat.of().formatHex(result));
}
}
MessageDigest.update(ByteBuffer)
consumes the bytes between the buffer’s position and limit. In buffered mode, flip() exposes only
the bytes just read, including a short final chunk, and clear() prepares the buffer for reuse.
The empty-file branch computes the digest without creating a mapping.
Create a three-byte input containing abc, with no trailing newline, then compile and run both
modes. The subshell’s noclobber setting refuses to overwrite an existing example.txt. Compilation
replaces FileChecksum.class in this working directory; checksum runs only read the input and print
to standard output.
(set -o noclobber; printf 'abc' > example.txt) &&
javac FileChecksum.java &&
java FileChecksum mapped example.txt &&
java FileChecksum buffered example.txt
Both commands should print this checksum:
ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad
To reuse an existing file, skip its creation and pass its path to either mode, quoting paths that
contain spaces. An empty file succeeds and prints
e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855.
Missing files, invalid arguments, and an oversized mapping raise an exception: Java writes the
diagnostic to standard error and exits with a nonzero status. The program does not print a checksum
when the read fails.
Performance considerations
Mapping has setup costs. The FileChannel documentation
notes that reading a small amount of data conventionally can cost less than creating a mapping.
A sequential checksum also spends time hashing, so a faster file-access method might make little
difference to the total runtime.
Use the same representative, unchanging file for both modes, within the mapped size limit. First check that the digests match. In Bash, these commands measure the whole command, including JVM startup, opening the file, mapping or buffer allocation, hashing, and closing the channel:
time java FileChecksum mapped example.txt &&
time java FileChecksum buffered example.txt
Replace example.txt in both commands with your representative file. The three-byte example is a
correctness check, not a useful speed comparison. Repeat the measurements, reverse the order, and
keep the JDK, JVM options, input, and machine load the same. Compare the spread of elapsed times,
not just the fastest run. Earlier reads can populate the operating system’s file cache, so report
first-run and repeated-run observations separately; a new JVM does not imply a cold file cache.
These timings answer how long this command takes. They do not predict the performance of a
long-running service or a random-access workload.
Best practices
Try-with-resources closes the channel, but closing a channel does not unmap its buffer. The mapping remains valid until the buffer is garbage-collected. In a long-running application, repeated mappings can therefore accumulate before reclamation. This short command ends its JVM after one file; it does not demonstrate deterministic unmapping inside a service.
Avoid equating low heap usage with low total memory usage. The mapping reserves virtual address space, and accessed file pages consume physical memory. This code avoids a full-file heap array, but the digest implementation may still copy chunks internally. Available address space and system resources can prevent mapping even below the API’s region limit.
Thread safety considerations
Keep the example single-threaded. Although file channels support concurrent use,
buffers require synchronization when shared between threads.
Even a read-only buffer has a mutable position, which digest.update() advances. Giving each task
its own buffer and digest avoids sharing those mutable objects; it does not protect the underlying
file from another writer.
Common use cases
For a one-pass checksum, buffered reading is a straightforward starting point and has no single-region size limit. Mapping is worth investigating when a parser repeatedly looks up records by byte offset in an unchanged file: a mapped buffer exposes indexed access without an explicit read for each lookup. Measure that access pattern separately before changing a working implementation. A checksum comparison alone cannot establish a speedup for it.
