# Resource Types for Data Flow Analysis in Code-Graph-RAG: A Complete Reference

> Explore ten resource types like ENV, FILE, NETWORK, DATABASE, and ENDPOINT for data flow analysis in Code-Graph-RAG. Understand program elements and external system interactions.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: api-reference
- Published: 2026-09-04

---

**Code-Graph-RAG defines ten canonical resource types—including ENV, FILE, NETWORK, DATABASE, and ENDPOINT—that annotate program elements using the `resource::<TYPE>::<QUALIFIER>` schema to model data flow between functions and external systems.**

Code-Graph-RAG (vitali87/code-graph-rag) is an open-source framework that constructs knowledge graphs from source code to power retrieval-augmented generation systems. Understanding the **resource types for data flow analysis in Code-Graph-RAG** is essential for tracing how functions interact with environment variables, file systems, network endpoints, and databases through labeled graph edges.

## The Canonical Resource Identifier Schema

Every resource in the analysis pipeline follows a strict naming convention encoded in `codec/schema.proto` (generated to [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)). The **Resource** protobuf message uses the format:

```

resource::<TYPE>::<QUALIFIER>

```

The **`<TYPE>`** component represents the resource category, while the **`<QUALIFIER>`** identifies the specific instance, such as [`/etc/app.conf`](https://github.com/vitali87/code-graph-rag/blob/main//etc/app.conf) for files or `HOME` for environment variables. This schema enables the graph builder to create precise **READS_FROM**, **WRITES_TO**, and **CALLS** edges between code entities and external resources.

## Complete List of Data Flow Resource Types

The codebase defines ten distinct resource types that capture every major I/O and communication primitive:

**ENV** — Captures environment variable dependencies. Example: `resource::ENV::HOME`

**STDOUT** — Represents standard output streams. Example: `resource::STDOUT::<dynamic>`

**STDIN** — Represents standard input streams. Example: `resource::STDIN::<dynamic>`

**STDERR** — Represents standard error streams. Example: `resource::STDERR::<dynamic>`

**FILE** — Tracks file system access including absolute and relative paths. Example: `resource::FILE::/etc/app.conf`

**DATABASE** — Models database connection strings and handles. Example: `resource::DATABASE::dsn`

**SOCKET** — Marks low-level socket endpoints. Example: `resource::SOCKET::example.com:80`

**NETWORK** — Captures high-level network URLs. Example: `resource::NETWORK::http://example.com/api`

**ENDPOINT** — Identifies service-oriented endpoints for RPC-style calls. Example: `resource::ENDPOINT::inv__1a2::ANY /inventory/reserve`

**RPC** — Labels remote procedure call identifiers at the method level. Example: `resource::RPC::UserService.GetUser`

## How Resource Types Drive Graph Construction

The analysis engine uses these type prefixes to generate semantic edges in the knowledge graph. When the system detects a function reading from an environment variable, it creates a **READS_FROM** edge connecting the function node to `resource::ENV::<varname>`. Similarly, write operations generate **WRITES_TO** edges, while service invocations create **CALLS** edges to **ENDPOINT** or **RPC** resources.

This type-aware labeling allows Code-Graph-RAG to answer complex queries about data provenance, such as tracing all functions that write to `resource::FILE::/tmp/sensitive.dat` or identifying code paths that read from `resource::ENV::API_KEY` before calling `resource::NETWORK::https://api.external.com`.

## Implementation Across the Test Suite

The repository validates resource type handling through comprehensive test files that demonstrate real-world usage patterns:

- **[`codebase_rag/tests/test_io_scala_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_io_scala_edges.py)** — Validates `ENV`, `FILE`, `STDOUT`, `STDERR`, and `NETWORK` resource edges in Scala codebases
- **[`codebase_rag/tests/test_io_libc_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_io_libc_edges.py)** — Demonstrates data flow analysis for standard I/O and environment variables in libc-dependent code
- **[`codebase_rag/tests/test_io_handle_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_io_handle_edges.py)** — Covers `FILE`, `DATABASE`, `SOCKET`, and `NETWORK` resource tracking for handle-based operations
- **[`codebase_rag/tests/test_route_call_endpoints.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_route_call_endpoints.py)** — Focuses on `ENDPOINT` and `NETWORK` resource resolution for service routing
- **[`codebase_rag/tests/unit/test_rust_cpp_flow_unit.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/unit/test_rust_cpp_flow_unit.py)** and **[`test_flat_flow_path_sensitive_unit.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_flat_flow_path_sensitive_unit.py)** — Verify path-sensitive flows from `ENV` to `STDOUT` in systems code

The underlying schema definition resides in `codec/schema.proto`, where the **Resource** message type establishes the enumeration of valid resource prefixes consumed by the graph builder.

## Practical Examples

The following patterns illustrate how resource identifiers appear in the analysis pipeline:

```python

# Environment variable access creates a READS_FROM edge to resource::ENV

env_path = "resource::ENV::PATH"

# File operations generate directed edges based on access mode

log_file = "resource::FILE::/var/log/app.log"

# Analysis result: function WRITE_TO -> log_file

# Network calls establish CALLS edges to NETWORK resources

api_endpoint = "resource::NETWORK::https://api.example.com/v1/data"

# Analysis result: function CALLS -> api_endpoint

```

In the test suite, assertions verify that code parsing correctly identifies these relationships:

```python

# From test_io_handle_edges.py: verifying database resource tracking

expected_resource = "resource::DATABASE::production_dsn"

# Validates that connection.open() READS_FROM expected_resource

```

## Summary

- Code-Graph-RAG uses a canonical `resource::<TYPE>::<QUALIFIER>` schema to annotate external dependencies
- **Ten resource types**—ENV, STDIN, STDOUT, STDERR, FILE, DATABASE, SOCKET, NETWORK, ENDPOINT, and RPC—cover all major I/O and communication patterns
- The graph builder creates **READS_FROM**, **WRITES_TO**, and **CALLS** edges based on these type annotations
- Implementation is validated across specialized test files including [`test_io_scala_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_io_scala_edges.py), [`test_io_handle_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_io_handle_edges.py), and [`test_route_call_endpoints.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_route_call_endpoints.py)
- The **Resource** protobuf definition in `codec/schema.proto` provides the authoritative type specification

## Frequently Asked Questions

### What is the format of a resource identifier in Code-Graph-RAG?

Resource identifiers follow the strict format `resource::<TYPE>::<QUALIFIER>`, where `<TYPE>` is one of ten defined categories (ENV, FILE, NETWORK, etc.) and `<QUALIFIER>` is the specific instance identifier such as a file path or variable name. This format is encoded in the **Resource** message defined in `codec/schema.proto`.

### How does Code-Graph-RAG use resource types for data flow analysis?

The analysis engine inspects function bodies to detect interactions with external systems. When it identifies a read operation from a file, it creates a **READS_FROM** edge from the function node to `resource::FILE::<path>`. Write operations generate **WRITES_TO** edges, while service calls create **CALLS** edges to **ENDPOINT** or **RPC** resources, enabling precise data lineage tracking.

### Where are the resource type definitions implemented in the repository?

The authoritative definitions reside in `codec/schema.proto` (compiled to [`codec/schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/codec/schema_pb2.py)), which specifies the **Resource** message structure. Usage examples and validation tests are distributed across [`codebase_rag/tests/test_io_scala_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/tests/test_io_scala_edges.py), [`test_io_libc_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_io_libc_edges.py), [`test_io_handle_edges.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_io_handle_edges.py), and [`test_route_call_endpoints.py`](https://github.com/vitali87/code-graph-rag/blob/main/test_route_call_endpoints.py).

### Can custom resource types be added to Code-Graph-RAG?

The current implementation uses the ten standard types defined in the protobuf schema: ENV, STDIN, STDOUT, STDERR, FILE, DATABASE, SOCKET, NETWORK, ENDPOINT, and RPC. To extend the system, you would need to modify the `codec/schema.proto` file to add new enum values to the **Resource** message and update the language-specific analyzers to detect and tag these new resource types.