How cavecrew-investigator Formats Output for Grep Parsing

The cavecrew-investigator agent outputs code symbol searches in a strict line-based format using path:line — \symbol` — note` with optional category headers (Defs:, Refs:, Callers:) to enable Unix-style text processing without JSON parsing.

The cavecrew-investigator is a read-only sub-agent in the JuliusBrussee/caveman repository designed specifically to locate code symbols and return them in a structured, grep-friendly format. Its output contract is defined in [agents/cavecrew-investigator.md](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md) and emphasizes predictable delimiters and minimal token usage for downstream LLM consumption.

The Line-Based Output Contract

Core Line Structure

Every result line follows a rigid pattern: <path:line> — \` — `.

  • <path:line> — The file path and line number separated by a colon (e.g., hooks/caveman-config.js:81).
  • \`` — The code identifier (function, variable, class, etc.) wrapped in backticks.
  • <note> — A concise description limited to six words or fewer.

For example:

hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW

Category Header Grouping

When three or more rows belong to the same category, the agent groups them under a one-word header followed by a colon. This makes it trivial for grep or other line-oriented tools to filter by category.

Common headers include:

  • Defs: — Definitions and declarations
  • Refs: — References and usages
  • Callers: — Call sites and invocations

Example output with headers:

Defs:
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW
- hooks/caveman-config.js:95 — `readFlag` — paired reader function

Refs:
- src/main.js:42 — `safeWriteFlag` — used in init

Edge Case Handling

Single-hit output: If only one match is found, the header is omitted and the line is printed directly:

hooks/caveman-config.js:160 — `readFlag` — paired reader

Zero-hit output: When nothing matches, the literal string No match. is emitted, which can be captured with grep -q to detect empty results.

Summary line: The final line optionally reports totals (e.g., 2 defs, 5 refs.). This line is omitted when the total is 0 or 1, keeping the output tidy for downstream parsers.

Practical Grep Parsing Examples

Because each line is plain text with a predictable delimiter (—), downstream scripts can filter for particular symbols, paths, or category headers without parsing JSON.

Filter for Definitions Only

caveman investigate "where is safeWriteFlag defined?" | grep '^Defs:' -A 10

Output:

Defs:
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW

Extract File Paths and Line Numbers

caveman investigate "list all uses of readFlag" | grep -oE '^[^ ]+:[0-9]+'

Output:

hooks/caveman-config.js:160

Count Callers Programmatically

caveman investigate "who calls activate?" | grep '^Callers:' -A 100 | wc -l

This returns the count including the header line, allowing automated analysis of dependency chains.

Design Philosophy

The format is deliberately caveman-compressed (as noted on line 6 of agents/cavecrew-investigator.md) to reduce token usage when results are fed back into the main LLM thread. By using plain text with standardized separators rather than JSON or XML, the agent maintains compatibility with traditional Unix text processing tools while remaining readable for both humans and language models.

The output contract is also documented in [skills/cavecrew/SKILL.md](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md) and referenced in [CLAUDE.md](https://github.com/JuliusBrussee/caveman/blob/main/CLAUDE.md) for Claude-based workflows, ensuring consistency across the caveman agent ecosystem.

Summary

  • Fixed format: Every line uses path:line — \symbol` — note` for predictable parsing.
  • Category headers: Defs, Refs, and Callers group results when three or more matches exist.
  • Edge cases: Single hits omit headers; zero hits return No match.; summary lines suppress for totals ≤ 1.
  • Unix-friendly: The em-dash (—) delimiter and plain text output enable standard grep, awk, and sed workflows without JSON overhead.
  • Token-optimized: The "caveman-compressed" design minimizes token usage for LLM consumption.

Frequently Asked Questions

What delimiter does cavecrew-investigator use between fields?

The agent uses an em-dash (—) surrounded by spaces to separate the location, symbol, and description fields. This character is distinct from standard hyphens, making it easy to target with grep or awk without conflicting with hyphenated file paths or symbol names.

How do I filter results to show only function definitions?

Pipe the output through grep '^Defs:' to capture only lines under the Definitions header. Since category headers are only emitted when three or more results exist, you may need to handle single-hit cases where the header is absent by checking for the presence of the backtick-wrapped symbol name instead.

What happens when cavecrew-investigator finds no matches?

The agent outputs the literal string No match. on a single line. You can detect this in shell scripts using grep -q "No match." or by checking if the output equals that exact string, allowing your automation to handle empty result sets gracefully.

Why doesn't cavecrew-investigator use JSON for its output?

According to the source code in agents/cavecrew-investigator.md, the format is intentionally "caveman-compressed" to reduce token usage when results are fed back into the main LLM thread. Plain text with predictable delimiters consumes fewer tokens than JSON syntax while remaining fully parseable by standard Unix text processing tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →