# How ripgrep Efficiently Applies Multiple Gitignore Glob Patterns

> Discover how ripgrep efficiently applies multiple gitignore glob patterns by compiling them into a single automaton for O(1) average time path matching with easy precedence rule resolution.

- Repository: [Andrew Gallant/ripgrep](https://github.com/BurntSushi/ripgrep)
- Tags: internals
- Published: 2026-03-05

---

**ripgrep compiles every glob pattern from all .gitignore files into a single deterministic automaton (GlobSet) that matches paths in O(1) average time, using reverse iteration to resolve precedence rules.**

When recursively searching large codebases, ripgrep must evaluate thousands of ignore patterns across nested directory structures without introducing noticeable overhead. The tool accomplishes this through a sophisticated compilation strategy implemented in the `ignore` crate, where multiple gitignore glob patterns are transformed into a unified, high-performance matching engine.

## Compiling Globs into a Unified GlobSet

The efficiency of ripgrep's gitignore handling stems from front-loading pattern compilation into a deterministic finite automaton. Rather than testing each file against individual glob strings, the system builds a `GlobSet` that can evaluate all patterns simultaneously.

### Parsing Individual Patterns

The process begins with `GitignoreBuilder::new(root)`, which initializes a builder for a specific directory scope. When `builder.add(path)` is called to ingest a .gitignore file, each line is processed by `GitignoreBuilder::add_line` (located in [`crates/ignore/src/gitignore.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/gitignore.rs) at lines 451-527).

This method transforms raw pattern strings into `Glob` structs that store:
- The original pattern string for reporting
- The compiled glob form with metadata (whitelist status, directory-only flags)
- The source file path for debugging

### Building the Deterministic Automaton

Once all gitignore files are added, `builder.build()` invokes `GlobSetBuilder::build()` (see [`crates/ignore/src/gitignore.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/gitignore.rs) lines 338-353). This creates a **single deterministic automaton** capable of testing a path against all registered globs in one pass, achieving O(1) average time complexity regardless of the number of patterns.

### The Gitignore Struct Internals

The resulting `Gitignore` struct (defined at lines 78-85 in [`crates/ignore/src/gitignore.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/gitignore.rs)) contains:
- `set: GlobSet` — the compiled matcher
- `globs: Vec<Glob>` — a parallel vector preserving insertion order for precedence resolution
- `matches: Option<Arc<Pool<Vec<usize>>>>` — a thread-local pool of reusable index buffers to eliminate per-match heap allocations

## Matching Paths Against Compiled Patterns

When `Gitignore::matched(path, is_dir)` evaluates a candidate file, it executes a highly optimized matching pipeline that minimizes string operations and memory allocations.

### Path Normalization and Stripping

First, `Gitignore::strip` (lines 73-99 in [`crates/ignore/src/gitignore.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/gitignore.rs)) normalizes the candidate path by removing:
- Leading `./` prefixes
- The global root directory (`self.root`)
- Leading slashes

This ensures the automaton receives a stable, relative string representation, avoiding repeated `Path` operations during matching.

### Vector-Indexed Matching with GlobSet

The method obtains a reusable buffer from the thread-local pool via `self.matches.as_ref().unwrap().get()`. It then wraps the stripped path in a `Candidate` and calls `self.set.matches_candidate_into(&candidate, &mut *matches)` (see `Gitignore::matched_stripped` at lines 57-65).

This fills the buffer with the **indices of all matching globs in registration order** (lowest index = first added), allowing the system to identify which patterns matched without re-running the automaton.

### Resolving Precedence via Reverse Iteration

Gitignore semantics dictate that later patterns override earlier ones. To resolve this efficiently, the code iterates the match buffer **in reverse** (`for &i in matches.iter().rev()`) as implemented in `Gitignore::matched_stripped` (lines 60-71).

During this reverse scan:
- The first non-directory-only glob (or directory-only match when `is_dir` is true) determines the result
- If `glob.is_whitelist()` returns true, the result is `Match::Whitelist`
- Otherwise, the result is `Match::Ignore`
- If the buffer is empty, `Match::None` is returned

This approach eliminates sorting overhead while correctly implementing gitignore precedence rules.

## Handling Multiple Gitignore Files During Directory Traversal

During filesystem traversal, ripgrep manages multiple .gitignore files through a layered architecture implemented in [`crates/ignore/src/dir.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/dir.rs) and [`crates/ignore/src/walk.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/walk.rs).

For each directory encountered, the system maintains a **stack of `Gitignore` objects** representing:
1. Global ignore patterns (e.g., `~/.config/git/ignore`)
2. The current directory's `.gitignore`
3. Optional per-directory overrides (`.ignore`, `.rgignore` from [`crates/ignore/src/overrides.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/overrides.rs))

The traversal logic in [`walk.rs`](https://github.com/BurntSushi/ripgrep/blob/main/walk.rs) (around line 727) consults this stack from most specific to most generic, using `Gitignore::matched_path_or_any_parents` to properly honor parent-directory patterns. Because each `Gitignore` instance already contains **all globs from its source file compiled into a single `GlobSet`**, checking an entry requires only a few cheap automaton traversals and a reverse scan of a small index vector.

## Performance Optimizations in ripgrep's Ignore Crate

The implementation leverages several specific optimizations to achieve negligible overhead even with thousands of patterns:

| Technique | Implementation Location | Performance Impact |
|-----------|------------------------|-------------------|
| **GlobSet Automaton** | `globset` crate via `GlobSetBuilder` | O(1) average matching regardless of pattern count |
| **Reverse Index Iteration** | `Gitignore::matched_stripped` (lines 60-71) | Eliminates sorting overhead for precedence resolution |
| **Thread-Local Buffer Pool** | `Arc<Pool<Vec<usize>>>` in `Gitignore` struct | Reuses match buffers, eliminating per-match heap allocations |
| **Path Stripping** | `Gitignore::strip` (lines 73-99) | Converts paths to stable relative strings before matching |
| **Early Empty Check** | `if self.is_empty()` short-circuit | Skips processing when no globs are present |
| **Literal Separator Optimization** | `GlobBuilder::literal_separator(true)` | Prevents `/**/` from crossing directory boundaries, keeping automaton size small |

These optimizations combine to allow ripgrep to evaluate complex ignore rules across massive repositories with minimal performance penalty.

## Practical Implementation Example

The following Rust code demonstrates the exact API that ripgrep uses internally to compile and match gitignore patterns:

```rust
use ignore::gitignore::{GitignoreBuilder, Match};
use std::path::Path;

// 1️⃣ Build a matcher from several .gitignore files
let mut builder = GitignoreBuilder::new("/my/project");
builder.add(Path::new(".gitignore")).unwrap();      // top-level
builder.add(Path::new("src/.gitignore")).unwrap();  // per-directory
let gitignore = builder.build().unwrap();

// 2️⃣ Test a path
let result = gitignore.matched("src/generated/tmp.rs", false);
match result {
    Match::Ignore(g) => println!("Ignored by: {}", g.original()),
    Match::Whitelist(g) => println!("Whitelisted by: {}", g.original()),
    Match::None => println!("File is not ignored"),
}

```

Internally, this creates a single `GlobSet` automaton containing all patterns from both `.gitignore` files. When `matched()` is called, ripgrep reuses a thread-local buffer to collect match indices, iterates them in reverse to resolve precedence, and returns the result without re-executing the automaton.

## Summary

- **Unified Compilation**: ripgrep compiles all gitignore patterns into a single `GlobSet` automaton via `GitignoreBuilder`, enabling O(1) average-time matching regardless of pattern count.
- **Efficient Precedence Resolution**: The system uses reverse iteration over match indices (`matches.iter().rev()`) to implement gitignore precedence rules without sorting overhead.
- **Memory Optimization**: Thread-local buffer pools (`Arc<Pool<Vec<usize>>>`) eliminate per-match heap allocations by reusing match index vectors.
- **Layered Architecture**: During traversal, ripgrep maintains a stack of compiled `Gitignore` objects (from [`crates/ignore/src/dir.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/dir.rs)) to handle multiple ignore files with minimal overhead.

## Frequently Asked Questions

### How does ripgrep handle conflicting gitignore patterns?

When multiple patterns match a path, ripgrep resolves conflicts by iterating match indices in reverse order (from `Gitignore::matched_stripped` in [`crates/ignore/src/gitignore.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/gitignore.rs)). Since later patterns in a .gitignore file have higher precedence according to git semantics, reverse iteration ensures the first applicable match encountered represents the correct rule. The system checks whether this winning pattern is a whitelist (negation) or standard ignore rule to determine the final result.

### What makes ripgrep's glob matching faster than checking patterns individually?

Rather than testing each pattern sequentially against a path, ripgrep compiles all globs into a single `GlobSet` automaton using the `globset` crate. This automaton processes the path in one pass with O(length_of_path) complexity, regardless of how many patterns are registered. Additionally, the implementation uses `literal_separator(true)` in `GlobBuilder` to prevent `/**/` from crossing directory boundaries unintentionally, which keeps the automaton size small and matching fast.

### How does ripgrep manage memory when checking thousands of files?

The `Gitignore` struct maintains a thread-local pool of reusable `Vec<usize>` buffers via `Arc<Pool<Vec<usize>>>`. When `matched()` is called, the system retrieves a pre-allocated buffer from this pool to store match indices, then returns it after use. This eliminates per-match heap allocations, which is critical when traversing repositories with millions of files. The pool is shared across threads via `Arc`, ensuring safe concurrent access without allocation bottlenecks.

### Can ripgrep handle nested .gitignore files in subdirectories?

Yes, ripgrep handles nested .gitignore files through a layered stack architecture implemented in [`crates/ignore/src/dir.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/dir.rs) and [`crates/ignore/src/walk.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/walk.rs). During directory traversal, the system builds a stack of `Gitignore` objects for each directory, including global ignores, the current directory's `.gitignore`, and optional override files (`.ignore`, `.rgignore`). When checking a file, ripgrep consults this stack from most specific to most generic, using `Gitignore::matched_path_or_any_parents` to properly honor parent-directory patterns while maintaining the O(1) matching performance of the compiled `GlobSet` automaton.