# How Doc2Graph Uses EasyOCR for Text Extraction in Custom Inference

> Discover how Doc2Graph uses EasyOCR to extract text for custom inference, converting detected text and polygons into node features for graph neural networks.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**Doc2Graph leverages EasyOCR as its dedicated text extraction engine during custom inference, transforming detected text regions and bounding polygons into node features that populate the graph neural network.**

When processing custom document images with the `andreagemelli/doc2graph` repository, the library converts raw pixels into structured graph representations through a pipeline that begins with optical character recognition. The `inference()` function initiates this workflow by invoking EasyOCR to detect and recognize text, which subsequently becomes the foundational data for graph construction and relationship prediction.

## The Custom Inference Entry Point

The inference workflow begins in **[`doc2graph/inference.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/inference.py)**, where the `inference()` function orchestrates the pipeline. For custom image inputs, the function instantiates a `GraphBuilder` and triggers graph construction by calling `gb.get_graph(paths, "CUSTOM")`.

When the `GraphBuilder` receives the `"CUSTOM"` data type flag, it routes image processing to the private method `__fromIMG()` inside **[`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py)**. This method serves as the primary integration point where EasyOCR extracts textual content from each input image before the library converts these detections into graph nodes.

## EasyOCR Integration in GraphBuilder

Inside `GraphBuilder.__fromIMG()`, Doc2Graph initializes an EasyOCR reader and processes each image path to extract text and spatial coordinates.

### Initializing the EasyOCR Reader

The code instantiates a fresh EasyOCR reader for the processing loop, configured specifically for English text recognition:

```python
reader = easyocr.Reader(["en"])
result = reader.readtext(path, paragraph=True)

```

The `Reader` is initialized with the `["en"]` language pack, and `readtext()` executes with `paragraph=True` to group words into logical text blocks. This returns a list of detections where each entry contains the polygon vertices and the recognized string.

### Converting Polygons to Axis-Aligned Bounding Boxes

EasyOCR returns text regions as arbitrary quadrilaterals, but Doc2Graph requires axis-aligned bounding boxes for graph node positioning. The code extracts the first and third polygon points to define the rectangle:

```python
box = [
    int(r[0][0][0]), int(r[0][0][1]),  # Top-left (x0, y0)

    int(r[0][2][0]), int(r[0][2][1])   # Bottom-right (x1, y1)

]
boxes.append(box)
texts.append(r[1])

```

This conversion generates standard `[x0, y0, x1, y1]` coordinates stored in the `boxes` list, while the corresponding recognized strings populate the `texts` list.

### Populating Node Features

After processing all detections, the method aggregates the OCR results into feature dictionaries that the graph construction engine consumes:

```python
features["boxs"] = boxes
features["texts"] = texts

```

These arrays become the **node attributes** for the resulting graph, where each text region represents a distinct node with spatial and textual features.

## Graph Construction from OCR Data

With `features["boxs"]` and `features["texts"]` populated, the `GraphBuilder` proceeds to wire nodes into edges based on the configured `edge_type` parameter (defined in **[`doc2graph/data/preprocessing.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/preprocessing.py)**). The library supports two connectivity modes:

- **`fully`**: Creates fully-connected graphs where every text node connects to every other node
- **`knn`**: Connects each node to its k-nearest spatial neighbors

The resulting `dgl.graph` objects encapsulate the document structure, with EasyOCR-derived bounding boxes serving as node coordinates and text strings as node labels. These graphs feed directly into the GNN model defined in **[`doc2graph/models/graphs.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/models/graphs.py)** for relationship prediction.

## Running Custom Inference: Complete Example

To execute custom inference on a batch of images, initialize the pipeline with pretrained weights and image paths:

```python
from doc2graph.inference import inference

# Load pretrained checkpoint (must exist in CHECKPOINTS directory)

weights = ["funsd-bert.pt"]

# Define image paths for processing

image_paths = [
    "samples/invoice1.png",
    "samples/invoice2.jpg",
]

# Execute inference (device=-1 for CPU, or specify GPU ID)

inference(weights, image_paths, device=-1)

```

Behind the scenes, this invocation triggers the EasyOCR extraction flow in `__fromIMG()`, builds the graph representations, and returns predictions visualized on the original images.

## Summary

- **EasyOCR is the exclusive text extraction engine** for custom image inference in Doc2Graph, invoked inside `GraphBuilder.__fromIMG()` in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py).
- **Bounding box conversion** transforms EasyOCR's polygon outputs into axis-aligned `[x0, y0, x1, y1]` coordinates stored in `features["boxs"]`.
- **Node feature population** maps recognized text strings to `features["texts"]`, establishing the textual attributes for each graph node.
- **Graph wiring** connects OCR-derived nodes using `fully` or `knn` edge strategies before feeding the structure to the GNN model.

## Frequently Asked Questions

### What OCR engine does Doc2Graph use for custom images?

Doc2Graph uses **EasyOCR** as its sole text extraction engine when processing custom images through the `inference()` function. The implementation specifically calls `easyocr.Reader(["en"])` inside [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) to detect and recognize text regions before converting them into graph nodes.

### How does Doc2Graph convert EasyOCR output into graph nodes?

The library converts EasyOCR's polygon-based detections into axis-aligned bounding boxes by extracting the first and third vertices of each quadrilateral to form `[x0, y0, x1, y1]` coordinates. These boxes populate `features["boxs"]` while the recognized text strings populate `features["texts"]`, creating the node attributes for the document graph.

### Can Doc2Graph extract text in languages other than English?

Currently, the EasyOCR reader is hardcoded to `["en"]` (English) in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py). However, the codebase structure suggests future multilingual support could be implemented by modifying the language list passed to `easyocr.Reader()` and ensuring corresponding model weights are available.

### Where is the text extraction logic located in the Doc2Graph codebase?

The core text extraction logic resides in **[`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py)** within the `GraphBuilder.__fromIMG()` method (lines 20-46). This method instantiates the EasyOCR reader, processes images, converts polygons to bounding boxes, and prepares the feature dictionaries used for graph construction.