Filter files in Ruby: a practical guide
Efficient file filtering is essential for everything from log processing to media management. Ruby provides robust tools for file selection that go beyond simple pattern matching. Let's explore practical techniques that scale from basic to advanced scenarios.
Requirements
The examples below are tested with Ruby 3.4.10 and mime-types 3.6.0, using the mime-types-data 3.2024.0806 registry. Use a maintained Ruby release for new projects. Only the MIME lookup example needs the mime-types gem. In an existing Bundler project, add:
# Add to your Gemfile
gem 'mime-types', '3.6.0'
Run bundle install, then execute your scripts with bundle exec ruby script.rb. Outside Bundler,
install the same version directly:
gem install mime-types -v 3.6.0
Save the Ruby examples in a .rb file and run it with ruby script.rb. Replace the example
directories with your own paths. Results are arrays or enumerators; use p to inspect them, such as
p recent_images after the complete FileFilter example.
The mime-types gem employs modified semantic versioning to track both API changes and registry data updates. For further details, refer to the mime-types documentation.
Core filtering methods
Ruby's standard library offers several immediate solutions for file filtering:
# Filter by extension
docs_dir = '/docs'
pdf_files = Dir.children(docs_dir).select do |f|
File.file?(File.join(docs_dir, f)) && File.extname(f) == '.pdf'
end
# Filter by size (1MB threshold)
large_files = Dir.glob('*').select do |f|
begin
File.file?(f) && File.size(f) > 1_000_000
rescue Errno::ENOENT, Errno::EACCES => e
warn "Error accessing #{f}: #{e.message}"
false
end
end
# Filter by modification time (last 24 hours)
recent_files = Dir.glob('*').select do |f|
begin
File.file?(f) && File.mtime(f) > (Time.now - 86400)
rescue Errno::ENOENT, Errno::EACCES => e
warn "Error accessing #{f}: #{e.message}"
false
end
end
These methods use the File class utilities for quick checks without loading file contents.
The extension check returns names relative to docs_dir and is case-sensitive. The glob examples
return paths relative to the current working directory and omit dotfiles. File.file? excludes
directories but follows symlinks to regular files; these checks are not a filesystem sandbox.
The time filters use a lower cutoff, so future-dated files also match.
Advanced pattern matching with glob
Ruby's Dir.glob supports UNIX-style pattern matching with some Ruby-specific enhancements:
# Match nested Markdown files
markdown_files = Dir.glob('**/*.md').select { |f| File.file?(f) }
# Match names containing a January 2024 date (not modification times)
jan_files = Dir.glob('*').grep(/(2024-01-\d{2})/).select { |f| File.file?(f) }
# Combined size and type filter
big_images = Dir.glob('*.{jpg,png}').select do |f|
begin
File.file?(f) && File.size(f) > 500_000
rescue Errno::ENOENT, Errno::EACCES => e
warn "Error accessing #{f}: #{e.message}"
false
end
end
Use double star (**) for recursive directory traversal and brace expansion for multiple
extensions.
MIME type lookup
Use the mime-types gem to look up the media type associated with a filename extension:
require 'mime/types'
def media_files(dir)
Dir.children(dir).select do |f|
begin
next unless File.file?(File.join(dir, f))
MIME::Types.type_for(f).any? do |mime|
mime.media_type == 'image' || mime.media_type == 'video'
end
rescue StandardError => e
warn "Error processing #{f}: #{e.message}"
false
end
end
end
# Usage:
visual_assets = media_files('/content/assets')
This approach uses an extension registry; it does not inspect file contents or validate that a file
is safe. A filename can have several registered types: .mp4, for example, has both application
and video matches in the tested registry. Checking every match avoids dropping it merely because
the first match is not a video type. It still does not prove that the file contains video.
Use content inspection separately when accepting untrusted uploads.
For additional details, see the
mime-types gem documentation.
Metadata filtering
Combine multiple metadata points for precise selection:
def recent_documents(path)
Dir.glob('*', base: path).map { |name| File.join(path, name) }.select do |f|
begin
next unless File.file?(f)
ext = File.extname(f).downcase
size = File.size(f)
modified = File.mtime(f)
(ext == '.pdf' || ext == '.docx') &&
size.between?(10_000, 5_000_000) &&
modified > (Time.now - 7*86400)
rescue StandardError => e
warn "Error processing #{f}: #{e.message}"
false
end
end
end
This selects PDF/DOCX files directly inside path, with case-insensitive extensions and sizes
from 10,000 through 5,000,000 bytes, inclusive. Files modified exactly seven days ago are excluded.
Passing the directory as base: keeps characters such as [ in its name out of the glob pattern.
Performance considerations
When processing large directories:
-
Lazy Evaluation: Defer metadata checks with
lazy.Dir.globstill builds its result array.Dir.glob('**/*').lazy .select { |f| File.file?(f) && File.size(f) > 1_000_000 } .first(10) -
Early Exit: Fail fast with
breakwhen possiblelog_dir = '/logs' Dir.children(log_dir).each do |name| f = File.join(log_dir, name) begin next unless File.file?(f) && f.end_with?('.log') break if File.size(f) > 1_000_000_000 # Stop at first huge log process_log(f) rescue StandardError => e warn "Error processing #{f}: #{e.message}" end endDefine
process_log(path)for your application's processing before running this excerpt. -
Metadata Caching: Store frequently accessed data
file_cache = {} Dir.glob('*').each do |f| begin file_cache[f] = { mtime: File.mtime(f), size: File.size(f) } rescue StandardError => e warn "Error caching #{f}: #{e.message}" end end
Production-grade example
This reusable selector returns an enumerator of paths under the supplied directory:
class FileFilter
def initialize(root_dir)
@root = root_dir
end
def find_files(extensions: [], min_size: 0, max_age: Float::INFINITY)
Dir.glob('**/*', base: @root).lazy.map { |name| File.join(@root, name) }.select do |path|
begin
next unless File.file?(path)
valid_extension = extensions.empty? || extensions.include?(File.extname(path))
valid_size = File.size(path) >= min_size
valid_age = (Time.now - File.mtime(path)) < max_age
valid_extension && valid_size && valid_age
rescue StandardError => e
warn "Error processing #{path}: #{e.message}"
false
end
end
end
end
# Usage:
filter = FileFilter.new('/user/uploads')
recent_images = filter.find_files(
extensions: ['.jpg', '.png'],
min_size: 100_000,
max_age: 3600 # 1 hour
).first(100)
The usage selects up to 100 .jpg or .png paths of at least 100,000 bytes, modified less than
one hour ago. Extension matching here is case-sensitive. Discovery still builds a glob array;
lazy defers the metadata checks, not traversal. Hidden files are omitted, symlinks to files are
followed, and symlinked directories are not recursively traversed. A missing root produces no
matches. Files can change after selection, so handle errors again when you open them.
Handling file encodings
Ruby typically handles file name encodings automatically when using UTF-8. However, if you encounter issues with non-ASCII characters in file names, create a UTF-8 display value as follows:
# Convert from Ruby's known source encoding and replace invalid bytes for display
utf8_labels = Dir.glob('*').map do |f|
f.encode('UTF-8', invalid: :replace, undef: :replace).scrub
end
Keep the original path for filesystem operations. Replacement characters can change a filename;
these labels are for display only. force_encoding merely relabels bytes and does not transcode.
Conclusion
Throughout this guide, we explored practical techniques to filter files in Ruby, from simple extension checks and glob patterns to filename-based MIME lookup and metadata filtering. By combining these methods, you can build efficient file processing pipelines tailored to your needs. For large-scale, production-ready solutions, consider exploring the robust file processing API offered by Transloadit. Visit our documentation to learn more.
