How ripgrep Efficiently Applies Multiple Gitignore Glob Patterns

ripgrep compiles every glob pattern from all .gitignore files into a single deterministic automaton (GlobSet) that matches paths in O(1) average time, using reverse iteration to resolve precedence rules.

When recursively searching large codebases, ripgrep must evaluate thousands of ignore patterns across nested directory structures without introducing noticeable overhead. The tool accomplishes this through a sophisticated compilation strategy implemented in the ignore crate, where multiple gitignore glob patterns are transformed into a unified, high-performance matching engine.

Compiling Globs into a Unified GlobSet

The efficiency of ripgrep's gitignore handling stems from front-loading pattern compilation into a deterministic finite automaton. Rather than testing each file against individual glob strings, the system builds a GlobSet that can evaluate all patterns simultaneously.

Parsing Individual Patterns

The process begins with GitignoreBuilder::new(root), which initializes a builder for a specific directory scope. When builder.add(path) is called to ingest a .gitignore file, each line is processed by GitignoreBuilder::add_line (located in crates/ignore/src/gitignore.rs at lines 451-527).

This method transforms raw pattern strings into Glob structs that store:

  • The original pattern string for reporting
  • The compiled glob form with metadata (whitelist status, directory-only flags)
  • The source file path for debugging

Building the Deterministic Automaton

Once all gitignore files are added, builder.build() invokes GlobSetBuilder::build() (see crates/ignore/src/gitignore.rs lines 338-353). This creates a single deterministic automaton capable of testing a path against all registered globs in one pass, achieving O(1) average time complexity regardless of the number of patterns.

The Gitignore Struct Internals

The resulting Gitignore struct (defined at lines 78-85 in crates/ignore/src/gitignore.rs) contains:

  • set: GlobSet — the compiled matcher
  • globs: Vec<Glob> — a parallel vector preserving insertion order for precedence resolution
  • matches: Option<Arc<Pool<Vec<usize>>>> — a thread-local pool of reusable index buffers to eliminate per-match heap allocations

Matching Paths Against Compiled Patterns

When Gitignore::matched(path, is_dir) evaluates a candidate file, it executes a highly optimized matching pipeline that minimizes string operations and memory allocations.

Path Normalization and Stripping

First, Gitignore::strip (lines 73-99 in crates/ignore/src/gitignore.rs) normalizes the candidate path by removing:

  • Leading ./ prefixes
  • The global root directory (self.root)
  • Leading slashes

This ensures the automaton receives a stable, relative string representation, avoiding repeated Path operations during matching.

Vector-Indexed Matching with GlobSet

The method obtains a reusable buffer from the thread-local pool via self.matches.as_ref().unwrap().get(). It then wraps the stripped path in a Candidate and calls self.set.matches_candidate_into(&candidate, &mut *matches) (see Gitignore::matched_stripped at lines 57-65).

This fills the buffer with the indices of all matching globs in registration order (lowest index = first added), allowing the system to identify which patterns matched without re-running the automaton.

Resolving Precedence via Reverse Iteration

Gitignore semantics dictate that later patterns override earlier ones. To resolve this efficiently, the code iterates the match buffer in reverse (for &i in matches.iter().rev()) as implemented in Gitignore::matched_stripped (lines 60-71).

During this reverse scan:

  • The first non-directory-only glob (or directory-only match when is_dir is true) determines the result
  • If glob.is_whitelist() returns true, the result is Match::Whitelist
  • Otherwise, the result is Match::Ignore
  • If the buffer is empty, Match::None is returned

This approach eliminates sorting overhead while correctly implementing gitignore precedence rules.

Handling Multiple Gitignore Files During Directory Traversal

During filesystem traversal, ripgrep manages multiple .gitignore files through a layered architecture implemented in crates/ignore/src/dir.rs and crates/ignore/src/walk.rs.

For each directory encountered, the system maintains a stack of Gitignore objects representing:

  1. Global ignore patterns (e.g., ~/.config/git/ignore)
  2. The current directory's .gitignore
  3. Optional per-directory overrides (.ignore, .rgignore from crates/ignore/src/overrides.rs)

The traversal logic in walk.rs (around line 727) consults this stack from most specific to most generic, using Gitignore::matched_path_or_any_parents to properly honor parent-directory patterns. Because each Gitignore instance already contains all globs from its source file compiled into a single GlobSet, checking an entry requires only a few cheap automaton traversals and a reverse scan of a small index vector.

Performance Optimizations in ripgrep's Ignore Crate

The implementation leverages several specific optimizations to achieve negligible overhead even with thousands of patterns:

Technique Implementation Location Performance Impact
GlobSet Automaton globset crate via GlobSetBuilder O(1) average matching regardless of pattern count
Reverse Index Iteration Gitignore::matched_stripped (lines 60-71) Eliminates sorting overhead for precedence resolution
Thread-Local Buffer Pool Arc<Pool<Vec<usize>>> in Gitignore struct Reuses match buffers, eliminating per-match heap allocations
Path Stripping Gitignore::strip (lines 73-99) Converts paths to stable relative strings before matching
Early Empty Check if self.is_empty() short-circuit Skips processing when no globs are present
Literal Separator Optimization GlobBuilder::literal_separator(true) Prevents /**/ from crossing directory boundaries, keeping automaton size small

These optimizations combine to allow ripgrep to evaluate complex ignore rules across massive repositories with minimal performance penalty.

Practical Implementation Example

The following Rust code demonstrates the exact API that ripgrep uses internally to compile and match gitignore patterns:

use ignore::gitignore::{GitignoreBuilder, Match};
use std::path::Path;

// 1️⃣ Build a matcher from several .gitignore files
let mut builder = GitignoreBuilder::new("/my/project");
builder.add(Path::new(".gitignore")).unwrap();      // top-level
builder.add(Path::new("src/.gitignore")).unwrap();  // per-directory
let gitignore = builder.build().unwrap();

// 2️⃣ Test a path
let result = gitignore.matched("src/generated/tmp.rs", false);
match result {
    Match::Ignore(g) => println!("Ignored by: {}", g.original()),
    Match::Whitelist(g) => println!("Whitelisted by: {}", g.original()),
    Match::None => println!("File is not ignored"),
}

Internally, this creates a single GlobSet automaton containing all patterns from both .gitignore files. When matched() is called, ripgrep reuses a thread-local buffer to collect match indices, iterates them in reverse to resolve precedence, and returns the result without re-executing the automaton.

Summary

  • Unified Compilation: ripgrep compiles all gitignore patterns into a single GlobSet automaton via GitignoreBuilder, enabling O(1) average-time matching regardless of pattern count.
  • Efficient Precedence Resolution: The system uses reverse iteration over match indices (matches.iter().rev()) to implement gitignore precedence rules without sorting overhead.
  • Memory Optimization: Thread-local buffer pools (Arc<Pool<Vec<usize>>>) eliminate per-match heap allocations by reusing match index vectors.
  • Layered Architecture: During traversal, ripgrep maintains a stack of compiled Gitignore objects (from crates/ignore/src/dir.rs) to handle multiple ignore files with minimal overhead.

Frequently Asked Questions

How does ripgrep handle conflicting gitignore patterns?

When multiple patterns match a path, ripgrep resolves conflicts by iterating match indices in reverse order (from Gitignore::matched_stripped in crates/ignore/src/gitignore.rs). Since later patterns in a .gitignore file have higher precedence according to git semantics, reverse iteration ensures the first applicable match encountered represents the correct rule. The system checks whether this winning pattern is a whitelist (negation) or standard ignore rule to determine the final result.

What makes ripgrep's glob matching faster than checking patterns individually?

Rather than testing each pattern sequentially against a path, ripgrep compiles all globs into a single GlobSet automaton using the globset crate. This automaton processes the path in one pass with O(length_of_path) complexity, regardless of how many patterns are registered. Additionally, the implementation uses literal_separator(true) in GlobBuilder to prevent /**/ from crossing directory boundaries unintentionally, which keeps the automaton size small and matching fast.

How does ripgrep manage memory when checking thousands of files?

The Gitignore struct maintains a thread-local pool of reusable Vec<usize> buffers via Arc<Pool<Vec<usize>>>. When matched() is called, the system retrieves a pre-allocated buffer from this pool to store match indices, then returns it after use. This eliminates per-match heap allocations, which is critical when traversing repositories with millions of files. The pool is shared across threads via Arc, ensuring safe concurrent access without allocation bottlenecks.

Can ripgrep handle nested .gitignore files in subdirectories?

Yes, ripgrep handles nested .gitignore files through a layered stack architecture implemented in crates/ignore/src/dir.rs and crates/ignore/src/walk.rs. During directory traversal, the system builds a stack of Gitignore objects for each directory, including global ignores, the current directory's .gitignore, and optional override files (.ignore, .rgignore). When checking a file, ripgrep consults this stack from most specific to most generic, using Gitignore::matched_path_or_any_parents to properly honor parent-directory patterns while maintaining the O(1) matching performance of the compiled GlobSet automaton.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →