# How the Memory Loop Persists Q&A Results for Re-Ingestion in code-review-graph

> Discover how the memory loop in code-review-graph persists Q&A results using markdown files for efficient re-ingestion and analysis runs.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: internals
- Published: 2026-08-18

---

**The memory loop in `tirth8205/code-review-graph` persists Q&A results as self-contained markdown files with YAML front-matter, enabling multiple analysis runs to build upon previous sessions without recomputing expensive graph queries.**

The `code-review-graph` project implements a lightweight persistence layer that captures interactive code‑review sessions as structured documents. By storing each question‑answer pair alongside its associated graph nodes, the system creates a cache that subsequent analyses can re‑ingest to accelerate feedback loops and maintain provenance across multiple runs.

## Core Persistence Functions in [`code_review_graph/memory.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/memory.py)

The memory subsystem centers on three public functions implemented in [[`code_review_graph/memory.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/memory.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/memory.py). Together they form a complete CRUD‑like interface for the Q&A cache.

### Saving Results with `save_result`

**`save_result`** writes a deterministic markdown file when a Q&A session concludes. The function signature accepts:

- `question` — the original user query
- `answer` — the generated response
- `nodes` — optional list of related graph node identifiers
- `result_type` — classification tag (e.g., `"query"`, `"analysis"`)
- `repo_root` — repository path where `.code-review-graph/memory/` resides

When only `repo_root` is supplied, the function automatically resolves the target directory to `<repo>/.code-review-graph/memory/`.

Each persisted file contains:

1. **YAML front‑matter** with `type`, `timestamp`, and the `nodes` list
2. **ATX heading** (`# <question>`) followed by the answer body

The filename is derived deterministically from the question text plus a timestamp, ensuring human‑readable ordering without collisions. See the implementation at lines 14‑75.

```python
from pathlib import Path
from code_review_graph.memory import save_result

repo_root = Path.cwd()
save_result(
    question="What functions handle memory?",
    answer="The memory module provides save_result, list_memories, and clear_memories.",
    nodes=["code_review_graph.memory.save_result", "code_review_graph.memory.list_memories"],
    result_type="query",
    repo_root=repo_root,
)

```

### Discovering Persisted Memories with `list_memories`

**`list_memories`** enables re‑ingestion by scanning the memory directory for `*.md` files, parsing their YAML front‑matter, and extracting the first heading as the original question. The function returns a list of dictionaries containing:

- `path` — absolute file location
- `question` — recovered query string
- `type` — result classification from front‑matter
- `timestamp` — ISO‑formatted creation time

This discoverability is essential for the re‑ingestion pipeline: downstream processes can enumerate prior sessions without prior knowledge of their contents. Implementation at lines 77‑118.

```python
from code_review_graph.memory import list_memories

memories = list_memories(repo_root=Path.cwd())
for m in memories:
    print(f"{m['timestamp']}: {m['question']} ({m.get('type')})")

```

### Cache Invalidation with `clear_memories`

**`clear_memories`** removes all markdown files in the memory directory and returns the deletion count. This utility supports test isolation, cache refresh workflows, and GDPR‑style data clearing scenarios. Implementation at lines 121‑141.

```python
from code_review_graph.memory import clear_memories

deleted = clear_memories(repo_root=Path.cwd())
print(f"Removed {deleted} memory files")

```

## How the Memory Loop Enables Re-Ingestion

The persistence architecture creates a bidirectional data flow:

```

        question
            │
            ▼
    ┌───────────────┐
    │  save_result  │──► markdown file (persisted)
    └───────────────┘      │
                           │
            ┌──────────────┘
            ▼
    ┌───────────────┐     ┌───────────────┐     ┌─────────────────┐
    │ list_memories │◄────│  read file    │────►│ feed into graph │
    └───────────────┘     └───────────────┘     └─────────────────┘

```

By decoupling **generation** from **consumption**, the system gains three capabilities:

1. **Incremental analysis** — New code‑review passes can layer atop previous Q&A context rather than rebuilding the graph from scratch.
2. **Provenance tracking** — Every result carries an immutable timestamp and node references for audit trails.
3. **Cross‑session learning** — An agent workflow described in [`AGENTS.md`](https://github.com/tirth8205/code-review-graph/blob/main/AGENTS.md) leverages persisted memories to refine recommendations over time.

## Key Files Supporting the Memory Subsystem

| File | Purpose |
|------|---------|
| [`code_review_graph/memory.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/memory.py) | Core persistence API: `save_result`, `list_memories`, `clear_memories` |
| [`tests/test_memory.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_memory.py) | Unit tests verifying correct markdown creation, parsing, and deletion |
| [`README.md`](https://github.com/tirth8205/code-review-graph/blob/main/README.md) | Architecture overview explaining the memory cache role |
| [`AGENTS.md`](https://github.com/tirth8205/code-review-graph/blob/main/AGENTS.md) | Agent workflow documentation describing continuous improvement via persisted memories |

## Summary

- **`save_result`** persists Q&A pairs as timestamped markdown files with YAML front‑matter in `.code-review-graph/memory/`.
- **`list_memories`** makes the cache discoverable by parsing filenames, front‑matter, and headings into structured dictionaries.
- **`clear_memories`** provides cache invalidation for testing and refresh workflows.
- The markdown‑on‑disk format enables **re‑ingestion** without re‑execution, supports **provenance tracking**, and powers **incremental agent learning** as described in project documentation.

## Frequently Asked Questions

### What format does the memory loop use for persistence?

The memory loop uses markdown files with YAML front‑matter. Each file stores metadata (`type`, `timestamp`, `nodes`) in the front‑matter block, followed by a level‑1 heading containing the original question and the answer body as prose.

### Where are memory files stored in the repository?

Memory files are automatically placed in `<repo_root>/.code-review-graph/memory/` when the `repo_root` parameter is supplied to `save_result`. The directory is created on first write if it does not exist.

### How does re‑ingestion differ from re‑running the original query?

Re‑ingestion reads persisted markdown files via `list_memories`, which is orders of magnitude faster than re‑executing the original graph traversal and LLM calls. The trade‑off is that re‑ingested results reflect the state at the time of original persistence, not live code changes.

### Can the memory cache be versioned or synced across machines?

The markdown‑file format is inherently portable and git‑friendly. Teams can commit the `.code-review-graph/memory/` directory to share institutional knowledge, though they should consider `.gitignore` policies for sensitive queries or large caches.