How the Memory Loop Persists Q&A Results for Re-Ingestion in code-review-graph

The memory loop in tirth8205/code-review-graph persists Q&A results as self-contained markdown files with YAML front-matter, enabling multiple analysis runs to build upon previous sessions without recomputing expensive graph queries.

The code-review-graph project implements a lightweight persistence layer that captures interactive code‑review sessions as structured documents. By storing each question‑answer pair alongside its associated graph nodes, the system creates a cache that subsequent analyses can re‑ingest to accelerate feedback loops and maintain provenance across multiple runs.

Core Persistence Functions in code_review_graph/memory.py

The memory subsystem centers on three public functions implemented in [code_review_graph/memory.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/memory.py). Together they form a complete CRUD‑like interface for the Q&A cache.

Saving Results with save_result

save_result writes a deterministic markdown file when a Q&A session concludes. The function signature accepts:

  • question — the original user query
  • answer — the generated response
  • nodes — optional list of related graph node identifiers
  • result_type — classification tag (e.g., "query", "analysis")
  • repo_root — repository path where .code-review-graph/memory/ resides

When only repo_root is supplied, the function automatically resolves the target directory to <repo>/.code-review-graph/memory/.

Each persisted file contains:

  1. YAML front‑matter with type, timestamp, and the nodes list
  2. ATX heading (# <question>) followed by the answer body

The filename is derived deterministically from the question text plus a timestamp, ensuring human‑readable ordering without collisions. See the implementation at lines 14‑75.

from pathlib import Path
from code_review_graph.memory import save_result

repo_root = Path.cwd()
save_result(
    question="What functions handle memory?",
    answer="The memory module provides save_result, list_memories, and clear_memories.",
    nodes=["code_review_graph.memory.save_result", "code_review_graph.memory.list_memories"],
    result_type="query",
    repo_root=repo_root,
)

Discovering Persisted Memories with list_memories

list_memories enables re‑ingestion by scanning the memory directory for *.md files, parsing their YAML front‑matter, and extracting the first heading as the original question. The function returns a list of dictionaries containing:

  • path — absolute file location
  • question — recovered query string
  • type — result classification from front‑matter
  • timestamp — ISO‑formatted creation time

This discoverability is essential for the re‑ingestion pipeline: downstream processes can enumerate prior sessions without prior knowledge of their contents. Implementation at lines 77‑118.

from code_review_graph.memory import list_memories

memories = list_memories(repo_root=Path.cwd())
for m in memories:
    print(f"{m['timestamp']}: {m['question']} ({m.get('type')})")

Cache Invalidation with clear_memories

clear_memories removes all markdown files in the memory directory and returns the deletion count. This utility supports test isolation, cache refresh workflows, and GDPR‑style data clearing scenarios. Implementation at lines 121‑141.

from code_review_graph.memory import clear_memories

deleted = clear_memories(repo_root=Path.cwd())
print(f"Removed {deleted} memory files")

How the Memory Loop Enables Re-Ingestion

The persistence architecture creates a bidirectional data flow:


        question
            │
            ▼
    ┌───────────────┐
    │  save_result  │──► markdown file (persisted)
    └───────────────┘      │
                           │
            ┌──────────────┘
            ▼
    ┌───────────────┐     ┌───────────────┐     ┌─────────────────┐
    │ list_memories │◄────│  read file    │────►│ feed into graph │
    └───────────────┘     └───────────────┘     └─────────────────┘

By decoupling generation from consumption, the system gains three capabilities:

  1. Incremental analysis — New code‑review passes can layer atop previous Q&A context rather than rebuilding the graph from scratch.
  2. Provenance tracking — Every result carries an immutable timestamp and node references for audit trails.
  3. Cross‑session learning — An agent workflow described in AGENTS.md leverages persisted memories to refine recommendations over time.

Key Files Supporting the Memory Subsystem

File Purpose
code_review_graph/memory.py Core persistence API: save_result, list_memories, clear_memories
tests/test_memory.py Unit tests verifying correct markdown creation, parsing, and deletion
README.md Architecture overview explaining the memory cache role
AGENTS.md Agent workflow documentation describing continuous improvement via persisted memories

Summary

  • save_result persists Q&A pairs as timestamped markdown files with YAML front‑matter in .code-review-graph/memory/.
  • list_memories makes the cache discoverable by parsing filenames, front‑matter, and headings into structured dictionaries.
  • clear_memories provides cache invalidation for testing and refresh workflows.
  • The markdown‑on‑disk format enables re‑ingestion without re‑execution, supports provenance tracking, and powers incremental agent learning as described in project documentation.

Frequently Asked Questions

What format does the memory loop use for persistence?

The memory loop uses markdown files with YAML front‑matter. Each file stores metadata (type, timestamp, nodes) in the front‑matter block, followed by a level‑1 heading containing the original question and the answer body as prose.

Where are memory files stored in the repository?

Memory files are automatically placed in <repo_root>/.code-review-graph/memory/ when the repo_root parameter is supplied to save_result. The directory is created on first write if it does not exist.

How does re‑ingestion differ from re‑running the original query?

Re‑ingestion reads persisted markdown files via list_memories, which is orders of magnitude faster than re‑executing the original graph traversal and LLM calls. The trade‑off is that re‑ingested results reflect the state at the time of original persistence, not live code changes.

Can the memory cache be versioned or synced across machines?

The markdown‑file format is inherently portable and git‑friendly. Teams can commit the .code-review-graph/memory/ directory to share institutional knowledge, though they should consider .gitignore policies for sensitive queries or large caches.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →