# How Code-Graph-RAG Handles Data Flow Analysis: Taint Tracking via Property Graphs

> Discover how Code-Graph-RAG handles data flow analysis using property graphs and FLOWS_TO edges for precise taint tracking. Learn to query value propagation with the FlowVerdict API.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: deep-dive
- Published: 2026-09-08

---

**Code-Graph-RAG performs data flow analysis by constructing a property graph with specialized FLOWS_TO edges that track value propagation from sources to sinks, then queries these paths using the FlowVerdict API to determine taint reachability.**

Code-Graph-RAG implements sophisticated **data flow analysis** by representing codebases as unified property graphs where every syntactic element becomes a node. The system captures value propagation through assignments, function calls, and I/O operations using dedicated **FLOWS_TO** relationships defined in [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py), enabling precise taint tracking across multiple programming languages. All flow relationships persist in Memgraph, allowing efficient path queries to determine if data can travel from a source to a potential sink.

## Graph Schema and FLOWS_TO Relationships

The foundation of data flow analysis in Code-Graph-RAG rests on a typed property graph schema. The `RelationshipType` enumeration in [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py) defines **FLOWS_TO** (line 952) as a first-class relationship type alongside standard edges like `CALLS`, `READS_FROM`, and `WRITES_TO`.

When the ingestion pipeline processes source files, it creates **FLOWS_TO** edges between expression nodes to represent value propagation. Each edge carries a lightweight "conduit" payload that records the specific mechanism of flow—whether via variable assignment, return values, or resource handles. This metadata enables precise tracking of how tainted data moves through the program structure.

## Constructing Flow Edges During Parsing

Language-specific parsers in the `codebase_rag/parsers/` directory—including [`type_inference.py`](https://github.com/vitali87/code-graph-rag/blob/main/type_inference.py) and [`structure_processor.py`](https://github.com/vitali87/code-graph-rag/blob/main/structure_processor.py)—walk the abstract syntax tree (AST) during the ingestion phase. When these traversers encounter assignments, function calls, or I/O operations, they emit **FLOWS_TO** edges connecting source expression nodes to destination expression nodes.

The parsing logic respects language-specific scoping rules, ensuring that flow edges accurately reflect lexical and dynamic visibility. For example, when a variable assignment occurs, the parser creates a unidirectional edge from the right-hand expression node to the left-hand variable node, marking the data dependency for subsequent analysis.

### Handling Complex Control Flow

Code-Graph-RAG's data flow engine accounts for advanced control flow scenarios that traditional static analysis might miss. The test suite validates branch-sensitive flow analysis in [`tests/test_flow_switch_branches.py`](https://github.com/vitali87/code-graph-rag/blob/main/tests/test_flow_switch_branches.py), ensuring that paths crossing conditional branches are tracked separately and only feasible flows are reported.

For resource-oriented programming patterns, the system handles handle-based writes as verified in [`tests/test_flow_handle_writes.py`](https://github.com/vitali87/code-graph-rag/blob/main/tests/test_flow_handle_writes.py). When a file handle or socket passes through multiple function calls, the taint propagates from the original source to the handle, then to the underlying resource, maintaining traceability even through abstraction layers.

## Querying Data Flows with the FlowVerdict API

The [`codebase_rag/flow_verdict.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/flow_verdict.py) module exposes the high-level **FlowVerdict** API, which answers the core question: "Can data flow from source X to sink Y?" This interface abstracts the underlying graph queries into a simple method call that accepts qualified names of functions or variables.

The verdict system returns three distinct states:
- **FOUND**: A directed path exists between the source and sink nodes
- **NOT_FOUND**: No path exists within the analyzed scope
- **UNKNOWN**: The analysis is incomplete due to missing type information or unresolvable symbols

### Cypher Query Implementation

Under the hood, `FlowVerdict` executes optimized Cypher queries against the Memgraph database. The query engine searches for directed paths composed exclusively of **FLOWS_TO** edges, or mixed paths combining `CALLS` and `FLOWS_TO` relationships where function call boundaries must be crossed.

```cypher
// Internal query pattern used by FlowVerdict
MATCH (src:Function {qualified_name: $source})
MATCH (sink:Function {qualified_name: $sink})
MATCH path = (src)-[:FLOWS_TO*]->(sink)
RETURN CASE WHEN path IS NULL THEN "NOT_FOUND" ELSE "FOUND" END AS verdict;

```

## Graph Service Architecture

The [`codebase_rag/services/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/graph_service.py) module provides the **GraphService** class, which manages connections to the Memgraph database using the Bolt protocol. This service handles node and edge persistence, ensuring that **FLOWS_TO** relationships remain consistent across ingestion batches.

Because the graph database stores flow relationships as native edges rather than adjacency lists, path queries execute with index-backed performance even on large codebases. The service layer abstracts connection management, transaction handling, and retry logic for the underlying Memgraph driver.

## Edge Maintenance and Cleanup

Long-running analysis sessions require maintenance of the flow graph's integrity. The [`codebase_rag/services/resource_cleanup.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/services/resource_cleanup.py) service periodically prunes dangling **FLOWS_TO** edges that reference deleted nodes or obsolete type information. This prevents false-positive flow paths where historical edges suggest connectivity that no longer exists in the current code state.

## Summary

- **FLOWS_TO edges** in [`types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/types_defs.py) provide the fundamental abstraction for tracking value propagation through assignments and calls.
- **Language parsers** emit these edges during AST traversal, capturing data dependencies at the expression level with conduit metadata.
- **FlowVerdict** in [`flow_verdict.py`](https://github.com/vitali87/code-graph-rag/blob/main/flow_verdict.py) offers a Python API to query reachability between sources and sinks, returning FOUND, NOT_FOUND, or UNKNOWN states.
- **Cypher path queries** against Memgraph enable fast traversal of flow chains, even across complex call graphs.
- **Branch-sensitive analysis** and handle-based resource tracking ensure accurate taint propagation through control flow and I/O abstractions.
- **Resource cleanup** services maintain graph integrity by removing stale flow edges during extended analysis sessions.

## Frequently Asked Questions

### How does Code-Graph-RAG represent data flow relationships?

Code-Graph-RAG represents data flow as native graph edges using the **FLOWS_TO** relationship type defined in [`codebase_rag/types_defs.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/types_defs.py). These edges connect source expression nodes to destination nodes, with optional payloads describing the conduit mechanism (variable, return value, or handle). The edges persist in Memgraph alongside structural relationships like `CALLS` and `CONTAINS`.

### What query language does Code-Graph-RAG use for flow analysis?

The system uses **Cypher**, the declarative query language native to Memgraph. The `FlowVerdict` class constructs Cypher queries that match variable-length paths using the `[:FLOWS_TO*]` syntax to find any number of hops between source and sink nodes. These queries execute against the Bolt protocol via the `GraphService` abstraction.

### How does the system handle language-specific parsing for flow edges?

Language-specific parsers in `codebase_rag/parsers/`—such as [`type_inference.py`](https://github.com/vitali87/code-graph-rag/blob/main/type_inference.py)—implement AST visitors that recognize assignment statements, return statements, and I/O operations. When these constructs are encountered, the parsers invoke graph mutation methods to create **FLOWS_TO** edges with appropriate source and target node references, respecting language-specific scoping and type resolution rules.

### What are the possible outcomes of a flow verdict query?

A `FlowVerdict.can_flow()` query returns three enum values: **FOUND** indicates a valid path exists via **FLOWS_TO** edges; **NOT_FOUND** confirms no path connects the specified qualified names; and **UNKNOWN** signals that the graph lacks sufficient type information or parsing coverage to determine connectivity definitively.