How to Perform Data Flow Analysis with Code-Graph-RAG: A Complete Guide

Use the flow_reachability_verdict() function in codebase_rag/flow_verdict.py to execute a breadth-first search across Neo4j FLOWS_TO relationships, returning a FlowVerdict named-tuple that reports whether data flows from a source qualified name to a sink.

Code-Graph-RAG (vitali87/code-graph-rag) is an open-source framework that maps codebases into a Neo4j graph database to enable intelligent static analysis. Performing data flow analysis with Code-Graph-RAG allows you to trace how tainted values propagate from source functions to sinks through the relationships stored in the graph.

How Data Flow Analysis Works in Code-Graph-RAG

The framework implements a flow-reachability verdict routine that operates in three distinct stages. It determines whether a data path exists between two symbols (qualified names) by querying the graph and performing an in-memory breadth-first search.

Stage 1: Collecting Flow Edges with Cypher

The analysis begins by extracting all FLOWS_TO edges that belong to the target project. The CYPHER_FLOW_EDGES query defined in codebase_rag/flow_verdict.py (lines 20–26) matches relationships where either endpoint belongs to the project:

MATCH (a)-[:FLOWS_TO]->(b)
WHERE a.qualified_name STARTS WITH $project_prefix
   OR b.qualified_name STARTS WITH $project_prefix
   OR a.qualified_name = $project_name
   OR b.qualified_name = $project_name
RETURN a.qualified_name AS source, b.qualified_name AS target

The graph_loader.py module executes this query via Neo4j sessions and returns rows as Python dictionaries, creating an adjacency list representation of the flow graph.

Stage 2: Breadth-First Search for Path Detection

Once edges are collected, the _bfs_path function (lines 97–120 in flow_verdict.py) performs the actual reachability analysis:

  • Graph representation: Edges are stored in a dict[str, list[str]] mapping source qualified names to their targets
  • Path reconstruction: The BFS builds a parent map to reconstruct the exact path when a sink is found
  • Cycle safety: A seen set prevents infinite loops by tracking visited nodes
  • Edge validation: The algorithm checks for the sink before checking the seen set, ensuring at least one real edge is traversed

This in-memory BFS operates on the adjacency list built from the Neo4j query results, avoiding expensive repeated graph traversals.

Stage 3: Handling Coverage Gaps and Uncertainty

If no path is found, the system runs the CYPHER_FLOW_COVERAGE_GAPS query (lines 33–40) to verify graph completeness:

  • UNKNOWN verdict: Returned when uncovered files exist (the analysis cannot guarantee absence of flow)
  • NO_FLOW verdict: Returned only when the analysis confirms the entire project is indexed with no possible paths

This distinction ensures the system avoids false negatives when static analysis lacks complete source coverage.

Implementing Data Flow Analysis in Python

To perform data flow analysis programmatically, initialize the GraphLoader and call the verdict function with qualified names (QNs) representing your source and sink.

Initializing the Analysis Pipeline

from codebase_rag.flow_verdict import flow_reachability_verdict
from codebase_rag.graph_loader import GraphLoader

# Create a connection to your Neo4j instance

loader = GraphLoader(uri="bolt://localhost:7687", auth=("neo4j", "password"))

# Define the project namespace and endpoint symbols

project = "my_project"
source_qn = "my_project.utils.input.read_user_input"
sink_qn = "my_project.db.write_to_db"

# Execute the analysis

verdict = flow_reachability_verdict(
    fetch_all=loader.fetch_all,
    project_name=project,
    source_qn=source_qn,
    sink_qn=sink_qn,
)

Interpreting the FlowVerdict Results

The flow_reachability_verdict function returns a FlowVerdict named-tuple with three possible states:

if verdict.verdict == "FOUND":
    # Path contains the sequence of qualified names

    print("Data flow detected:", " → ".join(verdict.path))
elif verdict.verdict == "UNKNOWN":
    # Gaps contains list of uncovered file paths

    print("Inconclusive analysis. Uncovered files:", verdict.gaps)
elif verdict.verdict == "NO_FLOW":
    print("Definitive: No data flow exists between endpoints")

Running Data Flow Checks from the Command Line

For ad-hoc analysis without writing Python scripts, use the CLI provided in codebase_rag/graph_cli.py:

python -m codebase_rag.graph_cli \
    --project my_project \
    --source my_project.utils.input.read_user_input \
    --sink my_project.db.write_to_db

The CLI forwards arguments directly to flow_reachability_verdict, making it suitable for CI/CD pipelines and security auditing workflows.

Summary

  • Flow-reachability verdict is the core mechanism for data flow analysis in Code-Graph-RAG, implemented in codebase_rag/flow_verdict.py
  • Three-stage process: Extract FLOWS_TO edges via Cypher, execute BFS in Python, and verify coverage gaps to determine UNKNOWN vs NO_FLOW states
  • GraphLoader (codebase_rag/graph_loader.py) manages Neo4j sessions and provides the fetch_all callback required by the verdict API
  • FlowVerdict returns "FOUND" with a path tuple, "NO_FLOW" for definitive negatives, or "UNKNOWN" with coverage gaps
  • CLI tool enables command-line execution via python -m codebase_rag.graph_cli

Frequently Asked Questions

What is the difference between NO_FLOW and UNKNOWN verdicts?

NO_FLOW indicates that the analysis has definitively confirmed no data path exists between the source and sink, requiring complete graph coverage of all project modules. UNKNOWN indicates that while no path was found in the existing graph, certain files lack flow-coverage metadata (stored in the gaps field), meaning a potential path might exist in unindexed code.

How does Code-Graph-RAG handle cycles during data flow analysis?

The _bfs_path function implements cycle detection using a seen set that tracks visited qualified names during traversal. According to the implementation in flow_verdict.py (lines 97–120), nodes are marked as seen after confirming they are not the target sink, ensuring the algorithm terminates correctly while still allowing the first edge to be traversed.

What are qualified names (QNs) in the context of Code-Graph-RAG?

Qualified names are unique identifiers stored in the qualified_name property of Neo4j nodes, typically following the format project.module.symbol (for example, my_project.utils.input.read_user_input). These strings serve as the primary keys for source and sink parameters in the flow_reachability_verdict function.

Can I perform data flow analysis without full graph coverage?

Yes, but the results will be conservative. If the graph lacks complete coverage of the target project, the system returns an UNKNOWN verdict rather than NO_FLOW, accompanied by a list of uncovered files in the gaps field. This design prevents false negatives in security analysis scenarios where untracked code might contain vulnerable data flows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →