How Doc2Graph Converts Document Images into Graph Representations

Doc2Graph converts document images into graph representations by extracting text and bounding boxes via OCR, constructing a DGL graph where nodes represent textual elements and edges encode spatial relationships, and enriching the graph with geometric, textual, and visual features for downstream processing.

The andreagemelli/doc2graph repository implements a modular pipeline that transforms scanned documents into rich graph structures suitable for graph neural networks. This process enables downstream tasks like key-value pair extraction by representing document layout and content as an interconnected graph rather than a flat text sequence.

Stage 1: OCR and Bounding-Box Extraction

The conversion process begins with optical character recognition to detect text regions and their spatial coordinates. Doc2Graph supports two primary OCR backends depending on the input source.

Tesseract OCR for Standard Inputs

When processing standard datasets or high-quality scans, the pipeline utilizes Tesseract via the pytesseract library. In doc2graph/data/preprocessing.py, the load_predictions function processes images and filters results by confidence scores:


# doc2graph/data/preprocessing.py

def load_predictions(path_preds, path_gts, path_images, debug=False):
    ...
    texts = pytesseract.image_to_data(Image.open(img_path), output_type=Output.DICT)
    for t in range(len(texts["level"])):
        if int(texts["conf"][t]) > 50 and texts["text"][t] != " ":
            b = [
                texts["left"][t],
                texts["top"][t],
                texts["left"][t] + texts["width"][t],
                texts["top"][t] + texts["height"][t],
            ]
            tp.append([b, texts["text"][t]])          # (bbox, text) pair

This function returns a list of bounding boxes (boxs) and associated OCR-extracted strings (texts), where each bounding box is defined by [x1, y1, x2, y2] coordinates.

EasyOCR for Custom Images

When processing custom image files through the CUSTOM mode, GraphBuilder.__fromIMG in doc2graph/data/graph_builder.py switches to EasyOCR for potentially better handling of diverse document layouts:


# doc2graph/data/graph_builder.py

reader = easyocr.Reader(["en"])
result = reader.readtext(path, paragraph=True)
for r in result:
    box = [
        int(r[0][0][0]), int(r[0][0][1]),
        int(r[0][2][0]), int(r[0][2][1]),
    ]
    boxs.append(box)
    texts.append(r[1])

Both methods produce standardized bounding box and text pairs that serve as the foundation for graph construction.

Stage 2: Graph Construction

With OCR data extracted, the pipeline constructs a dgl.DGLGraph object where document elements become nodes and spatial relationships become edges.

Node Creation from Bounding Boxes

In doc2graph/data/graph_builder.py, each detected bounding box generates a graph node. The implementation stores raw features for later enrichment:


# doc2graph/data/graph_builder.py (inside __fromIMG)

features["boxs"].append(boxs)   # store all bboxes

features["texts"].append(texts) # store all OCR strings

Node indices correspond directly to the position of each bounding box in the input list, creating a one-to-one mapping between spatial regions and graph vertices.

Edge Generation Strategies

Doc2Graph supports two configurable edge-creation strategies defined in doc2graph/data/graph_builder.py:

Fully-Connected Graphs

The fully_connected method creates complete graphs where every node connects to every other node:


# doc2graph/data/graph_builder.py

def fully_connected(self, ids: list) -> Tuple[list, list]:
    u, v = [], []
    for id in ids:
        u.extend([id for i in range(len(ids)) if i != id])
        v.extend([i for i in range(len(ids)) if i != id])
    return u, v

K-Nearest-Neighbour (KNN) Graphs

The knn_connection method creates sparser graphs by connecting each node only to its k spatially nearest neighbors:


# doc2graph/data/graph_builder.py

def knn_connection(self, size: tuple, bboxs: list, k=10) -> Tuple[list, list]:
    # 1. Build vertical / horizontal projection maps

    # 2. Expand a window around each bbox until ≥k neighbours are found

    # 3. Rank neighbours by polar distance (see utils.polar) and keep the nearest k

    return [e[0] for e in edges], [e[1] for e in edges]

The chosen strategy produces edge lists (u, v) that are passed to DGL to instantiate the graph structure:

g = dgl.graph((torch.tensor(u), torch.tensor(v)),
              num_nodes=len(boxs), idtype=torch.int32)

Stage 3: Feature Enrichment

After graph construction, FeatureBuilder.add_features in doc2graph/data/feature_builder.py injects rich node and edge attributes that combine geometric, textual, and visual information.

Node Feature Engineering

Each node receives a concatenated feature vector containing:

  • Geometric features (geom): Normalized bounding box coordinates stored in g.ndata["geom"]
  • Text histograms: Character distribution vectors generated via utils.get_histogram
  • Word embeddings: Dense vectors from SpaCy en_core_web_lg models
  • Visual embeddings: Features extracted through a ResNet-style backbone defined in doc2graph/models/unet/model.py

Edge Feature Engineering

Edges encode relational information through polar coordinate representations:


# doc2graph/data/feature_builder.py

dist, angle = polar(features["boxs"][id][pair[0]],
                    features["boxs"][id][pair[1]])
polar_coordinates = to_bin(distances, angles, self.num_polar_bins)
g.edata["feat"] = polar_coordinates

Additional edge attributes include:

  • Edge weights (g.edata["weights"]): Binary thresholds based on spatial distance (>0.9)
  • Node normalization (g.ndata["norm"]): Inverse degree of each node for message passing stabilization

End-to-End Inference Example

To convert document images into graph representations and extract key-value pairs, use the high-level inference function:

import torch
from doc2graph.inference import inference

# 1️⃣ Choose a checkpoint and the image(s) you want to process

weights = ["funsd-e2e-best.pt"]          # checkpoint name (must exist under `doc2graph/models/checkpoints`)

paths   = ["my_docs/doc1.png", "my_docs/doc2.png"]   # list of image files

# 2️⃣ Run inference – this builds graphs, enriches them and predicts links

inference(weights, paths, device=0)     # device=-1 for CPU

This execution flow performs:

  1. GraphBuilder().get_graph(paths, "CUSTOM") to generate base graphs and raw features
  2. FeatureBuilder().add_features(graphs, features) to inject geometric, textual, and visual embeddings
  3. Forward pass through the SetModel defined in doc2graph/models/graphs.py
  4. Output of visualizations to INFERENCE/<doc_name>.png and structured predictions to INFERENCE/<doc_name>.json

Summary

  • OCR Extraction: Doc2Graph uses Tesseract for standard inputs or EasyOCR for custom images to generate bounding boxes and text strings via doc2graph/data/preprocessing.py and doc2graph/data/graph_builder.py.
  • Graph Instantiation: Bounding boxes become nodes in a dgl.DGLGraph, with edges created through either fully-connected or K-nearest-neighbor strategies implemented in GraphBuilder.
  • Feature Enrichment: The FeatureBuilder class attaches geometric coordinates, SpaCy word embeddings, ResNet visual features, and polar-coordinate edge descriptors to prepare the graph for neural network processing.
  • Modular Configuration: Edge strategies and model parameters are controlled through YAML configuration files in configs/models/, allowing customization without code changes.

Frequently Asked Questions

What OCR engines does Doc2Graph support?

Doc2Graph supports both Tesseract and EasyOCR. The pipeline uses Tesseract via pytesseract.image_to_data() in doc2graph/data/preprocessing.py for standard dataset processing, while the __fromIMG method in doc2graph/data/graph_builder.py utilizes EasyOCR when processing custom images in CUSTOM mode.

How does Doc2Graph determine which document elements should be connected by edges?

Doc2Graph offers two edge-generation strategies configurable in the YAML files: fully-connected graphs where every node links to every other node via fully_connected(), and K-nearest-neighbour graphs where nodes connect only to their k closest spatial neighbors via knn_connection() in doc2graph/data/graph_builder.py.

What types of features are attached to nodes and edges in the graph?

Nodes receive geometric coordinates (normalized bounding boxes), text histograms, SpaCy word embeddings (en_core_web_lg), and visual embeddings from a ResNet backbone. Edges receive polar coordinate features (distance and angle binned via utils.polar and utils.to_bin) and distance-based weights, all added by FeatureBuilder.add_features() in doc2graph/data/feature_builder.py.

Can I process my own custom document images without retraining the model?

Yes. Place your images in a directory and call inference(weights, paths, device) from doc2graph/inference.py with a pretrained checkpoint. Set the mode to "CUSTOM" to trigger EasyOCR processing and graph construction for your specific documents, generating both visualization PNGs and structured JSON outputs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →