# How Doc2Graph Converts Document Images into Graph Representations

> Learn how Doc2Graph converts document images into graph representations by extracting text, detecting bounding boxes, and building a feature-rich DGL graph for advanced analysis.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**Doc2Graph converts document images into graph representations by extracting text and bounding boxes via OCR, constructing a DGL graph where nodes represent textual elements and edges encode spatial relationships, and enriching the graph with geometric, textual, and visual features for downstream processing.**

The `andreagemelli/doc2graph` repository implements a modular pipeline that transforms scanned documents into rich graph structures suitable for graph neural networks. This process enables downstream tasks like key-value pair extraction by representing document layout and content as an interconnected graph rather than a flat text sequence.

## Stage 1: OCR and Bounding-Box Extraction

The conversion process begins with optical character recognition to detect text regions and their spatial coordinates. Doc2Graph supports two primary OCR backends depending on the input source.

### Tesseract OCR for Standard Inputs

When processing standard datasets or high-quality scans, the pipeline utilizes Tesseract via the `pytesseract` library. In [`doc2graph/data/preprocessing.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/preprocessing.py), the `load_predictions` function processes images and filters results by confidence scores:

```python

# doc2graph/data/preprocessing.py

def load_predictions(path_preds, path_gts, path_images, debug=False):
    ...
    texts = pytesseract.image_to_data(Image.open(img_path), output_type=Output.DICT)
    for t in range(len(texts["level"])):
        if int(texts["conf"][t]) > 50 and texts["text"][t] != " ":
            b = [
                texts["left"][t],
                texts["top"][t],
                texts["left"][t] + texts["width"][t],
                texts["top"][t] + texts["height"][t],
            ]
            tp.append([b, texts["text"][t]])          # (bbox, text) pair

```

This function returns a list of bounding boxes (`boxs`) and associated OCR-extracted strings (`texts`), where each bounding box is defined by `[x1, y1, x2, y2]` coordinates.

### EasyOCR for Custom Images

When processing custom image files through the `CUSTOM` mode, `GraphBuilder.__fromIMG` in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) switches to **EasyOCR** for potentially better handling of diverse document layouts:

```python

# doc2graph/data/graph_builder.py

reader = easyocr.Reader(["en"])
result = reader.readtext(path, paragraph=True)
for r in result:
    box = [
        int(r[0][0][0]), int(r[0][0][1]),
        int(r[0][2][0]), int(r[0][2][1]),
    ]
    boxs.append(box)
    texts.append(r[1])

```

Both methods produce standardized bounding box and text pairs that serve as the foundation for graph construction.

## Stage 2: Graph Construction

With OCR data extracted, the pipeline constructs a `dgl.DGLGraph` object where document elements become nodes and spatial relationships become edges.

### Node Creation from Bounding Boxes

In [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py), each detected bounding box generates a graph node. The implementation stores raw features for later enrichment:

```python

# doc2graph/data/graph_builder.py (inside __fromIMG)

features["boxs"].append(boxs)   # store all bboxes

features["texts"].append(texts) # store all OCR strings

```

Node indices correspond directly to the position of each bounding box in the input list, creating a one-to-one mapping between spatial regions and graph vertices.

### Edge Generation Strategies

Doc2Graph supports two configurable edge-creation strategies defined in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py):

**Fully-Connected Graphs**

The `fully_connected` method creates complete graphs where every node connects to every other node:

```python

# doc2graph/data/graph_builder.py

def fully_connected(self, ids: list) -> Tuple[list, list]:
    u, v = [], []
    for id in ids:
        u.extend([id for i in range(len(ids)) if i != id])
        v.extend([i for i in range(len(ids)) if i != id])
    return u, v

```

**K-Nearest-Neighbour (KNN) Graphs**

The `knn_connection` method creates sparser graphs by connecting each node only to its *k* spatially nearest neighbors:

```python

# doc2graph/data/graph_builder.py

def knn_connection(self, size: tuple, bboxs: list, k=10) -> Tuple[list, list]:
    # 1. Build vertical / horizontal projection maps

    # 2. Expand a window around each bbox until ≥k neighbours are found

    # 3. Rank neighbours by polar distance (see utils.polar) and keep the nearest k

    return [e[0] for e in edges], [e[1] for e in edges]

```

The chosen strategy produces edge lists `(u, v)` that are passed to DGL to instantiate the graph structure:

```python
g = dgl.graph((torch.tensor(u), torch.tensor(v)),
              num_nodes=len(boxs), idtype=torch.int32)

```

## Stage 3: Feature Enrichment

After graph construction, `FeatureBuilder.add_features` in [`doc2graph/data/feature_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/feature_builder.py) injects rich node and edge attributes that combine geometric, textual, and visual information.

### Node Feature Engineering

Each node receives a concatenated feature vector containing:

- **Geometric features** (`geom`): Normalized bounding box coordinates stored in `g.ndata["geom"]`
- **Text histograms**: Character distribution vectors generated via `utils.get_histogram`
- **Word embeddings**: Dense vectors from SpaCy `en_core_web_lg` models
- **Visual embeddings**: Features extracted through a ResNet-style backbone defined in [`doc2graph/models/unet/model.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/unet/model.py)

### Edge Feature Engineering

Edges encode relational information through polar coordinate representations:

```python

# doc2graph/data/feature_builder.py

dist, angle = polar(features["boxs"][id][pair[0]],
                    features["boxs"][id][pair[1]])
polar_coordinates = to_bin(distances, angles, self.num_polar_bins)
g.edata["feat"] = polar_coordinates

```

Additional edge attributes include:
- **Edge weights** (`g.edata["weights"]`): Binary thresholds based on spatial distance (>0.9)
- **Node normalization** (`g.ndata["norm"]`): Inverse degree of each node for message passing stabilization

## End-to-End Inference Example

To convert document images into graph representations and extract key-value pairs, use the high-level `inference` function:

```python
import torch
from doc2graph.inference import inference

# 1️⃣ Choose a checkpoint and the image(s) you want to process

weights = ["funsd-e2e-best.pt"]          # checkpoint name (must exist under `doc2graph/models/checkpoints`)

paths   = ["my_docs/doc1.png", "my_docs/doc2.png"]   # list of image files

# 2️⃣ Run inference – this builds graphs, enriches them and predicts links

inference(weights, paths, device=0)     # device=-1 for CPU

```

This execution flow performs:
1. `GraphBuilder().get_graph(paths, "CUSTOM")` to generate base graphs and raw features
2. `FeatureBuilder().add_features(graphs, features)` to inject geometric, textual, and visual embeddings
3. Forward pass through the SetModel defined in [`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py)
4. Output of visualizations to `INFERENCE/<doc_name>.png` and structured predictions to `INFERENCE/<doc_name>.json`

## Summary

- **OCR Extraction**: Doc2Graph uses Tesseract for standard inputs or EasyOCR for custom images to generate bounding boxes and text strings via [`doc2graph/data/preprocessing.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/preprocessing.py) and [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py).
- **Graph Instantiation**: Bounding boxes become nodes in a `dgl.DGLGraph`, with edges created through either fully-connected or K-nearest-neighbor strategies implemented in `GraphBuilder`.
- **Feature Enrichment**: The `FeatureBuilder` class attaches geometric coordinates, SpaCy word embeddings, ResNet visual features, and polar-coordinate edge descriptors to prepare the graph for neural network processing.
- **Modular Configuration**: Edge strategies and model parameters are controlled through YAML configuration files in `configs/models/`, allowing customization without code changes.

## Frequently Asked Questions

### What OCR engines does Doc2Graph support?

Doc2Graph supports both Tesseract and EasyOCR. The pipeline uses Tesseract via `pytesseract.image_to_data()` in [`doc2graph/data/preprocessing.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/preprocessing.py) for standard dataset processing, while the `__fromIMG` method in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) utilizes EasyOCR when processing custom images in `CUSTOM` mode.

### How does Doc2Graph determine which document elements should be connected by edges?

Doc2Graph offers two edge-generation strategies configurable in the YAML files: **fully-connected** graphs where every node links to every other node via `fully_connected()`, and **K-nearest-neighbour** graphs where nodes connect only to their *k* closest spatial neighbors via `knn_connection()` in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py).

### What types of features are attached to nodes and edges in the graph?

Nodes receive geometric coordinates (normalized bounding boxes), text histograms, SpaCy word embeddings (`en_core_web_lg`), and visual embeddings from a ResNet backbone. Edges receive polar coordinate features (distance and angle binned via `utils.polar` and `utils.to_bin`) and distance-based weights, all added by `FeatureBuilder.add_features()` in [`doc2graph/data/feature_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/feature_builder.py).

### Can I process my own custom document images without retraining the model?

Yes. Place your images in a directory and call `inference(weights, paths, device)` from [`doc2graph/inference.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/inference.py) with a pretrained checkpoint. Set the mode to `"CUSTOM"` to trigger EasyOCR processing and graph construction for your specific documents, generating both visualization PNGs and structured JSON outputs.