How to Use ripgrep JSON Output Format for Programmatic Parsing

ripgrep's --json flag emits structured, line-delimited JSON records that enable robust programmatic parsing of search results without brittle text scraping.

The rg command in the BurntSushi/ripgrep repository supports a machine-readable output mode that serializes every match, context line, and file boundary into discrete JSON objects. This ripgrep JSON output format transforms the tool from a terminal utility into a programmable search engine component, streaming records that conform to a strict schema defined in the crates/printer source code.

Architecture of the JSON Printer

The JSON output implementation resides in the printer crate. In crates/printer/src/json.rs, the JSON struct acts as the top-level printer that receives internal message types and writes them as JSON Lines using serde_json.

The message protocol is defined in crates/printer/src/jsont.rs as a Message enum with four variants: Begin, End, Match, and Context. Each variant serializes to a JSON object containing a "type" field and a nested "data" object. The Data helper struct manages encoding, transparently handling both UTF-8 text and raw bytes via base64 to preserve non-UTF-8 content. The test suite in tests/json.rs demonstrates how to deserialize this output back into the same Rust types, validating the format's round-trip integrity.

Message Structure and Field Reference

Every file searched produces a deterministic sequence: a begin message, zero or more match or context messages, and a terminating end message. This ordering is guaranteed by the printer logic in crates/printer/src/json.rs.

The match type includes the following fields defined in the Match struct:

  • path: Optional file path as a string
  • lines: The matched text content
  • line_number: Optional line number
  • absolute_offset: Byte offset within the file
  • submatches: Array of SubMatch objects containing match (the text), start, and end byte offsets

The context type shares identical fields to match but represents lines provided for context via the -B, -A, or -C flags. The begin type signals the start of a new file search, while the end type delivers file-level statistics including searches_with_match, bytes_searched, and matched_lines.

Parsing ripgrep JSON Output in Practice

Filtering with jq in Shell Scripts

For command-line workflows, pipe rg --json to jq to extract specific fields without writing custom parsers:


# Extract only the matched text lines from all results

rg --json 'TODO' . | jq -r 'select(.type=="match") | .data.lines'

The select(.type=="match") filter ignores begin and end messages, while -r outputs raw strings without JSON quoting.

Stream Processing in Python

Python's built-in json module handles the JSON Lines format incrementally, maintaining constant memory usage for large searches:

#!/usr/bin/env python3
import json
import subprocess
import sys

proc = subprocess.Popen(
    ["rg", "--json", "TODO", "."],
    stdout=subprocess.PIPE,
    text=True,
)

for line in proc.stdout:
    msg = json.loads(line)
    if msg["type"] == "match":
        path = msg["data"]["path"]
        line_nr = msg["data"]["line_number"]
        text = msg["data"]["lines"]
        print(f"{path}:{line_nr}: {text}")

proc.wait()
sys.exit(proc.returncode)

This script spawns rg with the --json flag, iterates over the line-delimited output, and reconstructs a classic file:line:text format for each match object.

Type-Safe Deserialization in Rust

Consume the output using the same Message types defined in crates/printer/src/jsont.rs for compile-time schema validation:

use std::process::{Command, Stdio};
use serde::Deserialize;
use serde_json::Deserializer;

#[derive(Debug, Deserialize)]
#[serde(tag = "type", content = "data")]
enum Message {
    Begin { path: Option<String> },
    End { path: Option<String>, binary_offset: Option<u64>, stats: Stats },
    Match {
        path: Option<String>,
        lines: String,
        line_number: Option<u64>,
        absolute_offset: u64,
        submatches: Vec<SubMatch>,
    },
    Context { path: Option<String>, lines: String, line_number: Option<u64>, absolute_offset: u64, submatches: Vec<SubMatch> },
}

#[derive(Debug, Deserialize)]
struct Stats {
    searches_with_match: u64,
    bytes_searched: u64,
    bytes_printed: u64,
    matched_lines: u64,
    matches: u64,
}

#[derive(Debug, Deserialize)]
struct SubMatch {
    #[serde(rename = "match")]
    m: String,
    replacement: Option<String>,
    start: usize,
    end: usize,
}

fn main() {
    let mut child = Command::new("rg")
        .args(&["--json", "TODO", "."])
        .stdout(Stdio::piped())
        .spawn()
        .expect("failed to spawn rg");

    let stdout = child.stdout.take().expect("no stdout");
    let stream = Deserializer::from_reader(stdout).into_iter::<Message>();

    for msg in stream {
        match msg.expect("invalid json line") {
            Message::Match { path, line_number, lines, .. } => {
                println!("{}:{}: {}", path.unwrap_or_default(), line_number.unwrap_or(0), lines);
            }
            _ => {}
        }
    }

    let _ = child.wait();
}

The serde_json::Deserializer::from_reader approach streams the output without buffering the entire result set into memory, which is essential for searching large codebases.

Summary

  • ripgrep's --json flag produces JSON Lines output via the JSON printer implementation in crates/printer/src/json.rs
  • The Message enum in crates/printer/src/jsont.rs defines four serializable types: Begin, End, Match, and Context
  • Each line is a self-contained JSON object with a "type" discriminator and a "data" payload containing fields like path, lines, and submatches
  • Output follows a strict lifecycle per file: begin → match/context (zero or more) → end
  • The format supports incremental, streaming parsing suitable for integration into build pipelines and IDE plugins

Frequently Asked Questions

What is the difference between ripgrep's JSON and JSON Lines format?

ripgrep emits JSON Lines (also called NDJSON), where each line is an independent JSON object. This streaming format allows parsers to process results incrementally as they arrive via stdout, unlike a single monolithic JSON array that would require buffering the entire search output before parsing.

How do I handle non-UTF-8 data in ripgrep JSON output?

According to the Data helper implementation in crates/printer/src/jsont.rs, non-UTF-8 bytes are automatically base64-encoded to preserve binary data integrity. Your parser should inspect the data field for encoding indicators and decode base64 content when necessary to reconstruct the original bytes.

Can I parse ripgrep JSON output incrementally for large codebases?

Yes. The JSON Lines format is explicitly designed for streaming consumption. In Rust, use serde_json::Deserializer::from_reader; in Python, iterate over stdout line-by-line with json.loads(). This approach maintains constant memory usage regardless of the total result set size, as demonstrated in the test harness at tests/json.rs.

What fields are included in the "match" type JSON objects?

As defined in the Match struct in crates/printer/src/jsont.rs, match objects include path (optional string), lines (the matched text), line_number (optional integer), absolute_offset (byte position), and submatches (an array of SubMatch objects). Each SubMatch contains the exact match text with start and end byte offsets, enabling precise highlighting of patterns within a line.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →