What Post-Processing Optimizations Does the code-review-graph Pipeline Support?
The code-review-graph pipeline runs six deterministic post-processing steps after parsing: resolving call targets, computing node signatures, rebuilding the FTS5 search index, tracing execution flows, detecting code communities, and optionally refreshing embeddings.
The code_review_graph package enriches raw Tree-sitter parse results through a structured post-processing pipeline defined in [postprocessing.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py). These optimizations transform baseline syntax graphs into queryable, analysis-ready structures that power code search, navigation, and review workflows. Each step is non-fatal—failures are logged and aggregated without breaking the build.
How the Post-Processing Pipeline Works
The entry point run_post_processing orchestrates all optimizations sequentially. It accepts an optional embedding provider and model, wrapping each step in error handling that guarantees completion even when individual steps encounter issues.
The pipeline operates on a populated GraphStore instance (SQLite-backed) produced by either a full build or incremental update. Results are returned as a dictionary with per-step metrics, and warnings are collected in a shared list for diagnostic purposes.
Six Post-Processing Optimizations Explained
1. Resolve Bare and C++ Scoped Call Targets
Implementation: _resolve_bare_endpoints
This step populates "bare" edges—unresolved function calls that lack explicit targets—and resolves C++ scoped calls (Class::method syntax). The result is a graph where downstream analyses can follow concrete call relationships rather than ambiguous references.
2. Compute Node Signatures
Implementation: _compute_signatures
Generates human-readable signatures for nodes missing them, including functions, classes, and test definitions. These signatures enable quick identification in search results and UI displays without requiring full AST traversal.
3. Rebuild the FTS5 Full-Text Search Index
Implementation: _rebuild_fts_index
Recreates the SQLite FTS5 virtual table used for fast code search. This step ensures that newly parsed or updated code is immediately searchable through the package's query interface.
4. Trace Execution Flows
Implementation: _trace_flows
Walks the call graph from identified entry points, records flow edges, and persists them for flow-based queries. This enables "find all paths from X to Y" analyses and impact assessment features.
5. Detect Code Communities
Implementation: _detect_communities
Runs community detection using either the Leiden algorithm or simple file grouping, storing community IDs per node. Communities help identify modular boundaries and potential refactoring targets.
6. Refresh Embeddings (Optional)
Implementation: _refresh_embeddings
Only executes when embedding_provider and embedding_model parameters are explicitly supplied. Invokes the embedding service to update vector representations for semantic search capabilities.
Running the Pipeline: Code Examples
CLI Invocation
The code-review-graph command automatically triggers post-processing after parsing:
from code_review_graph.main import main_cli
# Equivalent to: code-review-graph run path/to/repo
main_cli()
Programmatic Use
Run the full pipeline from Python with optional embedding configuration:
from code_review_graph.graph import GraphStore
from code_review_graph.postprocessing import run_post_processing
with GraphStore("/tmp/graph.db") as store:
# Standard run without embeddings
results = run_post_processing(store)
# With embeddings enabled
results = run_post_processing(
store,
embedding_provider="openai",
embedding_model="text-embedding-ada-002",
)
print(results)
# Example output: {"bare_edges_resolved": 342, "signatures_computed": 156,
# "fts_indexed": 1245, "flows_traced": 89,
# "communities_detected": 12}
Selective Step Execution
Run individual optimizations for targeted updates:
from code_review_graph.graph import GraphStore
from code_review_graph.postprocessing import _compute_signatures
with GraphStore("graph.db") as store:
result = {}
warnings = []
_compute_signatures(store, result, warnings)
print(result) # {"signatures_computed": 57}
Key Source Files
| File | Purpose |
|---|---|
[postprocessing.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py) |
Core pipeline implementation with all six steps |
[graph.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) |
GraphStore class for SQLite graph operations |
[search.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) |
FTS5 index rebuilding utilities |
[flows.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/flows.py) |
Execution flow tracing and storage |
[communities.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py) |
Community detection algorithms |
[embeddings.py](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) |
Optional embedding refresh service |
Summary
- code-review-graph post-processing optimizations run deterministically after every parse, whether full or incremental builds
- The six core steps—call target resolution, signature computation, FTS5 index rebuild, flow tracing, community detection, and optional embedding refresh—enrich raw syntax graphs into analysis-ready structures
- All steps are non-fatal; the pipeline logs failures and returns partial results rather than breaking builds
- Configure embeddings explicitly via
embedding_providerandembedding_modelparameters - Access individual optimizations through underscored helper functions for selective updates
Frequently Asked Questions
Does the pipeline run automatically after parsing?
Yes. The CLI's run command and the programmatic build/incremental_update methods in graph.py both invoke run_post_processing immediately after the Tree-sitter parse completes. You do not need to call it manually unless running custom scripts against an existing database.
Can I disable specific post-processing steps?
There is no built-in configuration to skip steps in the standard pipeline. However, you can import individual functions from postprocessing.py (e.g., _compute_signatures, _trace_flows) and execute only the optimizations you need, as shown in the selective invocation example above.
What happens if the embedding refresh fails?
The embedding step is already optional—it only runs when both embedding_provider and embedding_model are provided. If the external service fails, the error is caught, logged, and added to the warnings list. The pipeline continues and returns results for all completed steps, following the same non-fatal behavior as other optimizations.
How do I check which optimizations ran successfully?
The run_post_processing function returns a dictionary with metric keys for each step: bare_edges_resolved, signatures_computed, fts_indexed, flows_traced, communities_detected, and optionally embeddings_refreshed. Check these keys to verify completion; examine the warnings list for any logged issues.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →