# How cavecrew-investigator Formats Output for Grep Parsing

> Learn how cavecrew-investigator formats output for grep parsing using a simple line-based structure. Process code symbol searches efficiently with Unix tools without complex JSON.

- Repository: [Julius Brussee/caveman](https://github.com/JuliusBrussee/caveman)
- Tags: how-to-guide
- Published: 2026-07-12

---

**The `cavecrew-investigator` agent outputs code symbol searches in a strict line-based format using `path:line — \`symbol\` — note` with optional category headers (Defs:, Refs:, Callers:) to enable Unix-style text processing without JSON parsing.**

The `cavecrew-investigator` is a read-only sub-agent in the [JuliusBrussee/caveman](https://github.com/JuliusBrussee/caveman) repository designed specifically to locate code symbols and return them in a structured, grep-friendly format. Its output contract is defined in [[`agents/cavecrew-investigator.md`](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md)](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md) and emphasizes predictable delimiters and minimal token usage for downstream LLM consumption.

## The Line-Based Output Contract

### Core Line Structure

Every result line follows a rigid pattern: `<path:line> — \`<symbol>\` — <note>`.

- **`<path:line>`** — The file path and line number separated by a colon (e.g., `hooks/caveman-config.js:81`).
- **`\`<symbol>\``** — The code identifier (function, variable, class, etc.) wrapped in backticks.
- **`<note>`** — A concise description limited to six words or fewer.

For example:

```text
hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW

```

### Category Header Grouping

When three or more rows belong to the same category, the agent groups them under a one-word header followed by a colon. This makes it trivial for `grep` or other line-oriented tools to filter by category.

Common headers include:

- **`Defs:`** — Definitions and declarations
- **`Refs:`** — References and usages
- **`Callers:`** — Call sites and invocations

Example output with headers:

```text
Defs:
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW
- hooks/caveman-config.js:95 — `readFlag` — paired reader function

Refs:
- src/main.js:42 — `safeWriteFlag` — used in init

```

### Edge Case Handling

**Single-hit output:** If only one match is found, the header is omitted and the line is printed directly:

```text
hooks/caveman-config.js:160 — `readFlag` — paired reader

```

**Zero-hit output:** When nothing matches, the literal string `No match.` is emitted, which can be captured with `grep -q` to detect empty results.

**Summary line:** The final line optionally reports totals (e.g., `2 defs, 5 refs.`). This line is omitted when the total is 0 or 1, keeping the output tidy for downstream parsers.

## Practical Grep Parsing Examples

Because each line is plain text with a predictable delimiter (`—`), downstream scripts can filter for particular symbols, paths, or category headers without parsing JSON.

### Filter for Definitions Only

```bash
caveman investigate "where is safeWriteFlag defined?" | grep '^Defs:' -A 10

```

Output:

```text
Defs:
- hooks/caveman-config.js:81 — `safeWriteFlag` — atomic write w/ O_NOFOLLOW

```

### Extract File Paths and Line Numbers

```bash
caveman investigate "list all uses of readFlag" | grep -oE '^[^ ]+:[0-9]+'

```

Output:

```text
hooks/caveman-config.js:160

```

### Count Callers Programmatically

```bash
caveman investigate "who calls activate?" | grep '^Callers:' -A 100 | wc -l

```

This returns the count including the header line, allowing automated analysis of dependency chains.

## Design Philosophy

The format is deliberately **caveman-compressed** (as noted on line 6 of [`agents/cavecrew-investigator.md`](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md)) to reduce token usage when results are fed back into the main LLM thread. By using plain text with standardized separators rather than JSON or XML, the agent maintains compatibility with traditional Unix text processing tools while remaining readable for both humans and language models.

The output contract is also documented in [[`skills/cavecrew/SKILL.md`](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md)](https://github.com/JuliusBrussee/caveman/blob/main/skills/cavecrew/SKILL.md) and referenced in [[`CLAUDE.md`](https://github.com/JuliusBrussee/caveman/blob/main/CLAUDE.md)](https://github.com/JuliusBrussee/caveman/blob/main/CLAUDE.md) for Claude-based workflows, ensuring consistency across the caveman agent ecosystem.

## Summary

- **Fixed format:** Every line uses `path:line — \`symbol\` — note` for predictable parsing.
- **Category headers:** Defs, Refs, and Callers group results when three or more matches exist.
- **Edge cases:** Single hits omit headers; zero hits return `No match.`; summary lines suppress for totals ≤ 1.
- **Unix-friendly:** The em-dash (`—`) delimiter and plain text output enable standard `grep`, `awk`, and `sed` workflows without JSON overhead.
- **Token-optimized:** The "caveman-compressed" design minimizes token usage for LLM consumption.

## Frequently Asked Questions

### What delimiter does cavecrew-investigator use between fields?

The agent uses an em-dash (`—`) surrounded by spaces to separate the location, symbol, and description fields. This character is distinct from standard hyphens, making it easy to target with `grep` or `awk` without conflicting with hyphenated file paths or symbol names.

### How do I filter results to show only function definitions?

Pipe the output through `grep '^Defs:'` to capture only lines under the Definitions header. Since category headers are only emitted when three or more results exist, you may need to handle single-hit cases where the header is absent by checking for the presence of the backtick-wrapped symbol name instead.

### What happens when cavecrew-investigator finds no matches?

The agent outputs the literal string `No match.` on a single line. You can detect this in shell scripts using `grep -q "No match."` or by checking if the output equals that exact string, allowing your automation to handle empty result sets gracefully.

### Why doesn't cavecrew-investigator use JSON for its output?

According to the source code in [`agents/cavecrew-investigator.md`](https://github.com/JuliusBrussee/caveman/blob/main/agents/cavecrew-investigator.md), the format is intentionally "caveman-compressed" to reduce token usage when results are fed back into the main LLM thread. Plain text with predictable delimiters consumes fewer tokens than JSON syntax while remaining fully parseable by standard Unix text processing tools.