How ripgrep Efficiently Applies Multiple Gitignore Glob Patterns
ripgrep compiles every glob pattern from all .gitignore files into a single deterministic automaton (GlobSet) that matches paths in O(1) average time, using reverse iteration to resolve precedence rules.
When recursively searching large codebases, ripgrep must evaluate thousands of ignore patterns across nested directory structures without introducing noticeable overhead. The tool accomplishes this through a sophisticated compilation strategy implemented in the ignore crate, where multiple gitignore glob patterns are transformed into a unified, high-performance matching engine.
Compiling Globs into a Unified GlobSet
The efficiency of ripgrep's gitignore handling stems from front-loading pattern compilation into a deterministic finite automaton. Rather than testing each file against individual glob strings, the system builds a GlobSet that can evaluate all patterns simultaneously.
Parsing Individual Patterns
The process begins with GitignoreBuilder::new(root), which initializes a builder for a specific directory scope. When builder.add(path) is called to ingest a .gitignore file, each line is processed by GitignoreBuilder::add_line (located in crates/ignore/src/gitignore.rs at lines 451-527).
This method transforms raw pattern strings into Glob structs that store:
- The original pattern string for reporting
- The compiled glob form with metadata (whitelist status, directory-only flags)
- The source file path for debugging
Building the Deterministic Automaton
Once all gitignore files are added, builder.build() invokes GlobSetBuilder::build() (see crates/ignore/src/gitignore.rs lines 338-353). This creates a single deterministic automaton capable of testing a path against all registered globs in one pass, achieving O(1) average time complexity regardless of the number of patterns.
The Gitignore Struct Internals
The resulting Gitignore struct (defined at lines 78-85 in crates/ignore/src/gitignore.rs) contains:
set: GlobSet— the compiled matcherglobs: Vec<Glob>— a parallel vector preserving insertion order for precedence resolutionmatches: Option<Arc<Pool<Vec<usize>>>>— a thread-local pool of reusable index buffers to eliminate per-match heap allocations
Matching Paths Against Compiled Patterns
When Gitignore::matched(path, is_dir) evaluates a candidate file, it executes a highly optimized matching pipeline that minimizes string operations and memory allocations.
Path Normalization and Stripping
First, Gitignore::strip (lines 73-99 in crates/ignore/src/gitignore.rs) normalizes the candidate path by removing:
- Leading
./prefixes - The global root directory (
self.root) - Leading slashes
This ensures the automaton receives a stable, relative string representation, avoiding repeated Path operations during matching.
Vector-Indexed Matching with GlobSet
The method obtains a reusable buffer from the thread-local pool via self.matches.as_ref().unwrap().get(). It then wraps the stripped path in a Candidate and calls self.set.matches_candidate_into(&candidate, &mut *matches) (see Gitignore::matched_stripped at lines 57-65).
This fills the buffer with the indices of all matching globs in registration order (lowest index = first added), allowing the system to identify which patterns matched without re-running the automaton.
Resolving Precedence via Reverse Iteration
Gitignore semantics dictate that later patterns override earlier ones. To resolve this efficiently, the code iterates the match buffer in reverse (for &i in matches.iter().rev()) as implemented in Gitignore::matched_stripped (lines 60-71).
During this reverse scan:
- The first non-directory-only glob (or directory-only match when
is_diris true) determines the result - If
glob.is_whitelist()returns true, the result isMatch::Whitelist - Otherwise, the result is
Match::Ignore - If the buffer is empty,
Match::Noneis returned
This approach eliminates sorting overhead while correctly implementing gitignore precedence rules.
Handling Multiple Gitignore Files During Directory Traversal
During filesystem traversal, ripgrep manages multiple .gitignore files through a layered architecture implemented in crates/ignore/src/dir.rs and crates/ignore/src/walk.rs.
For each directory encountered, the system maintains a stack of Gitignore objects representing:
- Global ignore patterns (e.g.,
~/.config/git/ignore) - The current directory's
.gitignore - Optional per-directory overrides (
.ignore,.rgignorefromcrates/ignore/src/overrides.rs)
The traversal logic in walk.rs (around line 727) consults this stack from most specific to most generic, using Gitignore::matched_path_or_any_parents to properly honor parent-directory patterns. Because each Gitignore instance already contains all globs from its source file compiled into a single GlobSet, checking an entry requires only a few cheap automaton traversals and a reverse scan of a small index vector.
Performance Optimizations in ripgrep's Ignore Crate
The implementation leverages several specific optimizations to achieve negligible overhead even with thousands of patterns:
| Technique | Implementation Location | Performance Impact |
|---|---|---|
| GlobSet Automaton | globset crate via GlobSetBuilder |
O(1) average matching regardless of pattern count |
| Reverse Index Iteration | Gitignore::matched_stripped (lines 60-71) |
Eliminates sorting overhead for precedence resolution |
| Thread-Local Buffer Pool | Arc<Pool<Vec<usize>>> in Gitignore struct |
Reuses match buffers, eliminating per-match heap allocations |
| Path Stripping | Gitignore::strip (lines 73-99) |
Converts paths to stable relative strings before matching |
| Early Empty Check | if self.is_empty() short-circuit |
Skips processing when no globs are present |
| Literal Separator Optimization | GlobBuilder::literal_separator(true) |
Prevents /**/ from crossing directory boundaries, keeping automaton size small |
These optimizations combine to allow ripgrep to evaluate complex ignore rules across massive repositories with minimal performance penalty.
Practical Implementation Example
The following Rust code demonstrates the exact API that ripgrep uses internally to compile and match gitignore patterns:
use ignore::gitignore::{GitignoreBuilder, Match};
use std::path::Path;
// 1️⃣ Build a matcher from several .gitignore files
let mut builder = GitignoreBuilder::new("/my/project");
builder.add(Path::new(".gitignore")).unwrap(); // top-level
builder.add(Path::new("src/.gitignore")).unwrap(); // per-directory
let gitignore = builder.build().unwrap();
// 2️⃣ Test a path
let result = gitignore.matched("src/generated/tmp.rs", false);
match result {
Match::Ignore(g) => println!("Ignored by: {}", g.original()),
Match::Whitelist(g) => println!("Whitelisted by: {}", g.original()),
Match::None => println!("File is not ignored"),
}
Internally, this creates a single GlobSet automaton containing all patterns from both .gitignore files. When matched() is called, ripgrep reuses a thread-local buffer to collect match indices, iterates them in reverse to resolve precedence, and returns the result without re-executing the automaton.
Summary
- Unified Compilation: ripgrep compiles all gitignore patterns into a single
GlobSetautomaton viaGitignoreBuilder, enabling O(1) average-time matching regardless of pattern count. - Efficient Precedence Resolution: The system uses reverse iteration over match indices (
matches.iter().rev()) to implement gitignore precedence rules without sorting overhead. - Memory Optimization: Thread-local buffer pools (
Arc<Pool<Vec<usize>>>) eliminate per-match heap allocations by reusing match index vectors. - Layered Architecture: During traversal, ripgrep maintains a stack of compiled
Gitignoreobjects (fromcrates/ignore/src/dir.rs) to handle multiple ignore files with minimal overhead.
Frequently Asked Questions
How does ripgrep handle conflicting gitignore patterns?
When multiple patterns match a path, ripgrep resolves conflicts by iterating match indices in reverse order (from Gitignore::matched_stripped in crates/ignore/src/gitignore.rs). Since later patterns in a .gitignore file have higher precedence according to git semantics, reverse iteration ensures the first applicable match encountered represents the correct rule. The system checks whether this winning pattern is a whitelist (negation) or standard ignore rule to determine the final result.
What makes ripgrep's glob matching faster than checking patterns individually?
Rather than testing each pattern sequentially against a path, ripgrep compiles all globs into a single GlobSet automaton using the globset crate. This automaton processes the path in one pass with O(length_of_path) complexity, regardless of how many patterns are registered. Additionally, the implementation uses literal_separator(true) in GlobBuilder to prevent /**/ from crossing directory boundaries unintentionally, which keeps the automaton size small and matching fast.
How does ripgrep manage memory when checking thousands of files?
The Gitignore struct maintains a thread-local pool of reusable Vec<usize> buffers via Arc<Pool<Vec<usize>>>. When matched() is called, the system retrieves a pre-allocated buffer from this pool to store match indices, then returns it after use. This eliminates per-match heap allocations, which is critical when traversing repositories with millions of files. The pool is shared across threads via Arc, ensuring safe concurrent access without allocation bottlenecks.
Can ripgrep handle nested .gitignore files in subdirectories?
Yes, ripgrep handles nested .gitignore files through a layered stack architecture implemented in crates/ignore/src/dir.rs and crates/ignore/src/walk.rs. During directory traversal, the system builds a stack of Gitignore objects for each directory, including global ignores, the current directory's .gitignore, and optional override files (.ignore, .rgignore). When checking a file, ripgrep consults this stack from most specific to most generic, using Gitignore::matched_path_or_any_parents to properly honor parent-directory patterns while maintaining the O(1) matching performance of the compiled GlobSet automaton.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →