# What Post-Processing Optimizations Does the code-review-graph Pipeline Support?

> Explore post-processing optimizations in the code-review-graph pipeline, including call target resolution, signature computation, FTS5 index rebuilding, execution flow tracing, community detection, and optional embedding refresh.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: performance
- Published: 2026-08-15

---

**The code-review-graph pipeline runs six deterministic post-processing steps after parsing: resolving call targets, computing node signatures, rebuilding the FTS5 search index, tracing execution flows, detecting code communities, and optionally refreshing embeddings.**

The `code_review_graph` package enriches raw Tree-sitter parse results through a structured post-processing pipeline defined in [[`postprocessing.py`](https://github.com/tirth8205/code-review-graph/blob/main/postprocessing.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py). These optimizations transform baseline syntax graphs into queryable, analysis-ready structures that power code search, navigation, and review workflows. Each step is non-fatal—failures are logged and aggregated without breaking the build.

## How the Post-Processing Pipeline Works

The entry point [`run_post_processing`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py) orchestrates all optimizations sequentially. It accepts an optional embedding provider and model, wrapping each step in error handling that guarantees completion even when individual steps encounter issues.

The pipeline operates on a populated `GraphStore` instance (SQLite-backed) produced by either a full build or incremental update. Results are returned as a dictionary with per-step metrics, and warnings are collected in a shared list for diagnostic purposes.

## Six Post-Processing Optimizations Explained

### 1. Resolve Bare and C++ Scoped Call Targets

**Implementation:** [`_resolve_bare_endpoints`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L73-L86)

This step populates "bare" edges—unresolved function calls that lack explicit targets—and resolves C++ scoped calls (`Class::method` syntax). The result is a graph where downstream analyses can follow concrete call relationships rather than ambiguous references.

### 2. Compute Node Signatures

**Implementation:** [`_compute_signatures`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L93-L122)

Generates human-readable signatures for nodes missing them, including functions, classes, and test definitions. These signatures enable quick identification in search results and UI displays without requiring full AST traversal.

### 3. Rebuild the FTS5 Full-Text Search Index

**Implementation:** [`_rebuild_fts_index`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L125-L138)

Recreates the SQLite FTS5 virtual table used for fast code search. This step ensures that newly parsed or updated code is immediately searchable through the package's query interface.

### 4. Trace Execution Flows

**Implementation:** [`_trace_flows`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L141-L155)

Walks the call graph from identified entry points, records flow edges, and persists them for flow-based queries. This enables "find all paths from X to Y" analyses and impact assessment features.

### 5. Detect Code Communities

**Implementation:** [`_detect_communities`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L158-L172)

Runs community detection using either the Leiden algorithm or simple file grouping, storing community IDs per node. Communities help identify modular boundaries and potential refactoring targets.

### 6. Refresh Embeddings (Optional)

**Implementation:** [`_refresh_embeddings`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py#L175-L203)

Only executes when `embedding_provider` and `embedding_model` parameters are explicitly supplied. Invokes the embedding service to update vector representations for semantic search capabilities.

## Running the Pipeline: Code Examples

### CLI Invocation

The `code-review-graph` command automatically triggers post-processing after parsing:

```python
from code_review_graph.main import main_cli

# Equivalent to: code-review-graph run path/to/repo

main_cli()

```

### Programmatic Use

Run the full pipeline from Python with optional embedding configuration:

```python
from code_review_graph.graph import GraphStore
from code_review_graph.postprocessing import run_post_processing

with GraphStore("/tmp/graph.db") as store:
    # Standard run without embeddings

    results = run_post_processing(store)
    
    # With embeddings enabled

    results = run_post_processing(
        store,
        embedding_provider="openai",
        embedding_model="text-embedding-ada-002",
    )
    
    print(results)
    # Example output: {"bare_edges_resolved": 342, "signatures_computed": 156, 

    #                  "fts_indexed": 1245, "flows_traced": 89, 

    #                  "communities_detected": 12}

```

### Selective Step Execution

Run individual optimizations for targeted updates:

```python
from code_review_graph.graph import GraphStore
from code_review_graph.postprocessing import _compute_signatures

with GraphStore("graph.db") as store:
    result = {}
    warnings = []
    _compute_signatures(store, result, warnings)
    
    print(result)  # {"signatures_computed": 57}

```

## Key Source Files

| File | Purpose |
|------|---------|
| [[`postprocessing.py`](https://github.com/tirth8205/code-review-graph/blob/main/postprocessing.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/postprocessing.py) | Core pipeline implementation with all six steps |
| [[`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/graph.py) | `GraphStore` class for SQLite graph operations |
| [[`search.py`](https://github.com/tirth8205/code-review-graph/blob/main/search.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/search.py) | FTS5 index rebuilding utilities |
| [[`flows.py`](https://github.com/tirth8205/code-review-graph/blob/main/flows.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/flows.py) | Execution flow tracing and storage |
| [[`communities.py`](https://github.com/tirth8205/code-review-graph/blob/main/communities.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/communities.py) | Community detection algorithms |
| [[`embeddings.py`](https://github.com/tirth8205/code-review-graph/blob/main/embeddings.py)](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/embeddings.py) | Optional embedding refresh service |

## Summary

- **code-review-graph post-processing optimizations** run deterministically after every parse, whether full or incremental builds
- The six core steps—call target resolution, signature computation, FTS5 index rebuild, flow tracing, community detection, and optional embedding refresh—enrich raw syntax graphs into analysis-ready structures
- All steps are **non-fatal**; the pipeline logs failures and returns partial results rather than breaking builds
- Configure embeddings explicitly via `embedding_provider` and `embedding_model` parameters
- Access individual optimizations through underscored helper functions for selective updates

## Frequently Asked Questions

### Does the pipeline run automatically after parsing?

Yes. The CLI's `run` command and the programmatic `build`/`incremental_update` methods in [`graph.py`](https://github.com/tirth8205/code-review-graph/blob/main/graph.py) both invoke `run_post_processing` immediately after the Tree-sitter parse completes. You do not need to call it manually unless running custom scripts against an existing database.

### Can I disable specific post-processing steps?

There is no built-in configuration to skip steps in the standard pipeline. However, you can import individual functions from [`postprocessing.py`](https://github.com/tirth8205/code-review-graph/blob/main/postprocessing.py) (e.g., `_compute_signatures`, `_trace_flows`) and execute only the optimizations you need, as shown in the selective invocation example above.

### What happens if the embedding refresh fails?

The embedding step is already optional—it only runs when both `embedding_provider` and `embedding_model` are provided. If the external service fails, the error is caught, logged, and added to the warnings list. The pipeline continues and returns results for all completed steps, following the same non-fatal behavior as other optimizations.

### How do I check which optimizations ran successfully?

The `run_post_processing` function returns a dictionary with metric keys for each step: `bare_edges_resolved`, `signatures_computed`, `fts_indexed`, `flows_traced`, `communities_detected`, and optionally `embeddings_refreshed`. Check these keys to verify completion; examine the `warnings` list for any logged issues.