# DGL RPC Vulnerabilities: Security Considerations for the doc2graph Repository

> Secure your doc2graph projects by understanding DGL RPC vulnerabilities. Learn how this repository mitigates risks by restricting Deep Graph Library operations locally, preventing unauthorized remote access.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: security
- Published: 2026-02-24

---

**The doc2graph repository eliminates DGL RPC attack surfaces by restricting all Deep Graph Library operations to local in-process graph construction and message-passing, never importing or invoking the `dgl.rpc` or `dgl.distributed` subsystems.**

When evaluating security considerations regarding DGL RPC vulnerabilities in document-to-graph conversion pipelines, it is critical to distinguish between local graph manipulation and distributed training architectures. The doc2graph project, hosted at `andreagemelli/doc2graph`, implements graph neural networks using DGL strictly for single-machine workloads, ensuring that remote code execution risks, unauthorized node access, and man-in-the-middle tampering threats inherent in DGL's RPC subsystem are not present in the current codebase.

## Local DGL Usage in doc2graph (No RPC Exposure)

The repository's architecture explicitly avoids distributed computing primitives. In **[`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py)**, the code imports DGL solely for local graph construction:

```python
import dgl

```

According to the source code at lines 57-89, this file handles graph building and manipulation entirely within the local Python process. Similarly, **[`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py)** imports only message-passing kernels:

```python
import dgl.function as fn

```

These imports at lines 3-4 support local neural network operations without network I/O. The entry points in **[`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py)** and **[`doc2graph/inference.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/inference.py)** invoke these modules but never enable RPC listeners or client connections. Because the codebase never calls `dgl.rpc` or `dgl.distributed` APIs, attack vectors such as malicious payload deserialization over network sockets or unauthorized RPC command execution are physically absent from the deployment surface.

## Potential DGL RPC Vulnerabilities in Distributed Extensions

While doc2graph currently operates safely in local mode, integrating DGL RPC for multi-machine scaling introduces specific security concerns that require mitigation.

### Untrusted Graph Deserialization

DGL RPC endpoints can deserialize graph data received over the network. A crafted payload could trigger arbitrary Python code execution on worker nodes, resulting in full **remote code execution (RCE)** compromise.

**Mitigation:** Validate and whitelist graph schemas before deserialization, use cryptographically signed or encoded payloads, and execute workers within restricted containers or sandboxed environments.

### Unauthenticated RPC Endpoints

By default, DGL opens TCP listeners without mandatory authentication mechanisms. Unauthorized parties could issue arbitrary RPC calls to trigger model inference, manipulate training states, or write to file systems.

**Mitigation:** Deploy TLS-encrypted communication channels or VPNs, enforce mutual authentication via certificates or token-based auth, and restrict network access to known peer nodes.

### Denial-of-Service (DoS) via Message Flooding

Flooding RPC messages can exhaust memory and CPU resources on distributed training servers, causing service outages or indefinite training stalls.

**Mitigation:** Implement rate-limiting on inbound RPC calls, monitor resource usage metrics, and apply back-pressure mechanisms to throttle excessive requests.

### Information Leakage Through Metadata

RPC metadata containing node identifiers, edge types, or graph structures may reveal sensitive document architecture or confidential relationships during transmission.

**Mitigation:** Encrypt all RPC payloads in transit, strip or anonymize sensitive attributes before serialization, and audit metadata fields for accidental data exposure.

### Binary Version Mismatch Exploits

Different DGL versions across cluster nodes may interpret binary graph data inconsistently, potentially causing crashes or silent data corruption that could destabilize training pipelines.

**Mitigation:** Pin identical DGL versions across all distributed nodes and perform compatibility checks at cluster startup.

## Code Examples: Safe Local Usage vs. Secure RPC Patterns

### Current Safe Local Implementation

The following pattern from doc2graph demonstrates secure local operation with zero network exposure:

```python
from doc2graph.data.graph_builder import GraphBuilder

gb = GraphBuilder()
graph, nodes, edges, feats = gb.get_graph(src_path="data/funsd", src_data="FUNSD")

# All operations happen in-process; no network traffic.

```

In **[`doc2graph/training/utils.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/utils.py)**, the `SetModel` helper (lines 35-47) selects model classes like `GCN`, `EDGE`, or `E2E` without invoking any distributed initialization routines.

### Secure RPC Wrapper for Future Distributed Scaling

If extending doc2graph to support distributed training, implement TLS encryption and explicit function registration:

```python
import dgl
import ssl
from dgl.rpc import Server, Client

# 1️⃣ Create a TLS context (self-signed certs for demo)

ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH)
ctx.load_cert_chain(certfile="cert.pem", keyfile="key.pem")

# 2️⃣ Start a server that only accepts connections from trusted peers

server = Server("0.0.0.0:50051", ssl_context=ctx)
server.register_function("process_graph", lambda g: g.ndata["feat"] * 2)
server.start()

# 3️⃣ Client side – authenticate using the same TLS context

client = Client("server-host:50051", ssl_context=ctx)
g = dgl.graph(([0, 1], [1, 2]))
processed = client.call("process_graph", g)   # Safe RPC call

```

Key security measures in this example include **TLS encryption** for transport security, **explicit allowlisting** of RPC functions, and validation of graph objects before processing.

## Summary

- **doc2graph uses DGL locally only**: The repository imports `dgl` and `dgl.function` strictly for in-process graph construction in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) and message-passing in [`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py), never invoking `dgl.rpc`.

- **No current attack surface**: Because the code never initializes RPC servers or clients, vulnerabilities such as remote code execution via graph deserialization or unauthorized endpoint access do not apply to the existing codebase.

- **Future RPC integration requires hardening**: If implementing distributed training later, address deserialization risks with schema validation, enable TLS mutual authentication, deploy rate-limiting for DoS protection, and encrypt metadata to prevent information leakage.

- **Version consistency matters**: Pin DGL versions across all nodes to prevent binary incompatibility issues that could cause crashes or data corruption in distributed settings.

## Frequently Asked Questions

### Does doc2graph currently use DGL RPC for distributed training?

No. According to the source code analysis, doc2graph never imports `dgl.rpc` or `dgl.distributed`. All graph operations in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) and [`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py) execute locally within a single Python process, eliminating network-based attack vectors.

### Can DGL RPC vulnerabilities affect single-machine inference with doc2graph?

No. DGL RPC vulnerabilities require active RPC server or client initialization to expose TCP endpoints or process remote messages. Since doc2graph's inference pipeline in [`doc2graph/inference.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/inference.py) operates entirely through local function calls without network serialization, RPC-specific threats such as man-in-the-middle attacks or payload injection are not applicable.

### How should I secure DGL RPC endpoints when scaling doc2graph to multiple machines?

Enable TLS encryption for all transport layers, implement mutual authentication using certificates or tokens, and register only explicitly allowed functions via `server.register_function()`. Additionally, run worker processes in sandboxed containers, validate graph schemas before deserialization, and apply rate-limiting to prevent DoS attacks.

### What specific DGL functions does doc2graph use for graph construction?

The repository uses standard DGL graph construction APIs in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) (lines 57-89) and message-passing functions via `import dgl.function as fn` in [`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py) (lines 3-4). These APIs handle local `DGLGraph` object manipulation and neural network message aggregation without distributed primitives.