How Ripgrep Memory-Mapped File Search Differs From Buffered Search
Ripgrep memory-mapped file search leverages the memmap crate to create a zero-copy view of file contents in virtual memory, while buffered search incrementally reads data into a reusable 8 KB LineBuffer, with each strategy selected by the Searcher type based on platform constraints and user configuration.
The BurntSushi/ripgrep repository implements two distinct file ingestion strategies within its Searcher architecture. Understanding how ripgrep memory-mapped file search contrasts with its buffered alternative reveals critical performance, safety, and compatibility trade-offs that affect search behavior across different operating systems and file types.
Core Implementation Architectures
Ripgrep’s search dispatcher, search_file_maybe_path in crates/searcher/src/searcher/mod.rs, routes file processing through two mutually exclusive code paths depending on whether memory-mapping succeeds.
Memory-Mapped Search Path
When memory-mapped file search is enabled and available, ripgrep invokes self.config.mmap.open(file, path) from crates/searcher/src/searcher/mmap.rs. If this returns Some(mmap), the searcher immediately delegates to search_slice, treating the mapped region as a &[u8] slice.
The implementation uses the memmap crate to create a read-only memory mapping. This approach eliminates user-space buffer copies because the kernel serves file pages directly to the pattern matching engine. The same line-by-line matching logic used for in-memory slices processes the mapped data without additional heap allocations for file content.
Buffered (Roll) Search Path
If the memory map attempt returns None, the dispatcher falls back to search_reader, which constructs a LineBufferReader defined in crates/searcher/src/line_buffer.rs. This buffered search strategy repeatedly fills a fixed-size LineBuffer (defaulting to 8 KB) using the ReadByLine implementation in crates/searcher/src/searcher/glue.rs.
Data streams from the file descriptor in chunks, with the buffer holding only a sliding window of the file at any moment. This approach requires explicit read syscalls and memory copies into the user-space buffer, but maintains constant memory footprint regardless of file size.
Activation Conditions and Platform Constraints
The choice between ripgrep memory-mapped file search and buffered search depends on the MmapChoice configuration and runtime platform detection.
MmapChoice Behavior:
MmapChoice::never()– Default setting that forces buffered search for all filesMmapChoice::auto()– Enables heuristic memory-mapping on supported platforms
Platform-Specific Disabling:
Even when auto() is selected, the open method in crates/searcher/src/searcher/mmap.rs explicitly disables memory-mapping on macOS (cfg!(target_os = "macos")). The maintainers observed poorer performance characteristics on macOS systems, forcing a fallback to buffered search regardless of file type.
File Type Limitations: Memory-mapping requires regular files with seekable descriptors. Pipes, stdin, and other special files automatically trigger the buffered path even when mapping is globally enabled.
Performance, Memory, and Safety Trade-offs
Understanding the operational differences between these strategies helps optimize search workloads.
Memory Usage Patterns:
- Memory-mapped: The entire file size becomes addressable in virtual memory space, though no additional heap allocation occurs beyond mapping metadata. Physical memory consumption depends on which pages the kernel caches.
- Buffered: Strictly bounded to the
LineBuffercapacity (8 KB by default), though the buffer may grow if a user-specifiedheap_limitrequires larger allocations for extremely long lines.
Speed Characteristics:
- Memory-mapped: Delivers superior throughput for large, frequently accessed files already resident in the page cache, avoiding per-read syscall overhead. However, the initial
mmapsyscall and page-fault handling create overhead that makes this approach slower for tiny files. - Buffered: Provides consistent performance across all file sizes and platforms, though it incurs read-and-copy overhead for every buffer fill operation.
Safety Considerations:
Memory-mapped file search carries a significant risk: if the underlying file is truncated while the map is active, accessing mapped pages beyond the new file boundary triggers a SIGBUS signal that aborts the process. This behavior is documented in the memory_map builder comments within the source. Buffered search faces no such hazard because all reads occur through explicit file descriptor operations with standard OS error handling.
Multi-Line Search Implications
Both strategies converge when handling multi-line patterns. Even when memory-mapping is enabled, multi-line searches require the entire file content to be accessible simultaneously. In memory-mapped file search, the map naturally provides this. For buffered search, rigprip automatically allocates a heap buffer large enough to hold the complete file contents, effectively negating the memory advantage of the rolling buffer for multi-line regex patterns.
Configuring Search Strategies in Code
Developers using ripgrep as a library can explicitly control the search strategy through the SearcherBuilder API:
use grep_searcher::SearcherBuilder;
use grep_searcher::searcher::MmapChoice;
// Default buffered search (MmapChoice::never)
let buffered_searcher = SearcherBuilder::new().build();
// Enable heuristic memory-mapped file search
let mmap_searcher = SearcherBuilder::new()
.memory_map(MmapChoice::auto())
.build();
When using MmapChoice::auto(), the searcher attempts to map files on Linux and Windows (excluding macOS), falling back automatically to LineBufferReader when mapping fails or is inappropriate for the file type.
Summary
- Ripgrep memory-mapped file search uses
memmap::Mmapviacrates/searcher/src/searcher/mmap.rsto create zero-copy file views processed bysearch_slice, while buffered search usesLineBufferandReadByLineincrates/searcher/src/searcher/glue.rsfor incremental processing. - Memory-mapping is disabled by default (
MmapChoice::never) and explicitly blocked on macOS due to performance degradation, requiring the buffered path. - Memory-mapped searches risk SIGBUS termination if files truncate during access, whereas buffered searches operate safely through file descriptors.
- Multi-line searches force both strategies to hold the complete file in memory, eliminating the buffered approach's typical memory advantage.
- The dispatcher in
crates/searcher/src/searcher/mod.rshandles strategy selection automatically based on platform, file type, and user configuration.
Frequently Asked Questions
Why does ripgrep disable memory-mapped file search on macOS?
According to the source code in crates/searcher/src/searcher/mmap.rs, ripgrep explicitly checks cfg!(target_os = "macos") and returns None from the open method, forcing a fallback to buffered search. The maintainers observed poorer performance characteristics on macOS systems compared to Linux and Windows, making memory-mapping counterproductive on that platform.
Can memory-mapped file searches crash ripgrep?
Yes. If a file is truncated while ripgrep holds a memory map, subsequent access to mapped pages beyond the new file boundary generates a SIGBUS signal that terminates the process immediately. This risk is documented in the memory_map builder comments and represents the primary safety disadvantage compared to buffered search, which handles truncated files gracefully through standard read error semantics.
How does ripgrep handle multi-line regex with buffered search?
When processing multi-line patterns, ripgrep requires random access to the entire haystack. Even in buffered mode, the Searcher allocates a heap buffer large enough to contain the complete file contents, effectively loading the entire file into memory regardless of the 8 KB rolling buffer configuration. This behavior ensures pattern matching correctness but increases memory consumption for large files during multi-line searches.
What is the default buffer size for ripgrep's buffered search?
The default LineBuffer capacity is 8 KB, defined in crates/searcher/src/line_buffer.rs. This buffer size balances syscall overhead against memory footprint, though the buffer can grow dynamically if a user specifies a heap_limit that requires accommodating lines longer than the default capacity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →