How Doc2Graph Uses EasyOCR for Text Extraction in Custom Inference
Doc2Graph leverages EasyOCR as its dedicated text extraction engine during custom inference, transforming detected text regions and bounding polygons into node features that populate the graph neural network.
When processing custom document images with the andreagemelli/doc2graph repository, the library converts raw pixels into structured graph representations through a pipeline that begins with optical character recognition. The inference() function initiates this workflow by invoking EasyOCR to detect and recognize text, which subsequently becomes the foundational data for graph construction and relationship prediction.
The Custom Inference Entry Point
The inference workflow begins in doc2graph/inference.py, where the inference() function orchestrates the pipeline. For custom image inputs, the function instantiates a GraphBuilder and triggers graph construction by calling gb.get_graph(paths, "CUSTOM").
When the GraphBuilder receives the "CUSTOM" data type flag, it routes image processing to the private method __fromIMG() inside doc2graph/data/graph_builder.py. This method serves as the primary integration point where EasyOCR extracts textual content from each input image before the library converts these detections into graph nodes.
EasyOCR Integration in GraphBuilder
Inside GraphBuilder.__fromIMG(), Doc2Graph initializes an EasyOCR reader and processes each image path to extract text and spatial coordinates.
Initializing the EasyOCR Reader
The code instantiates a fresh EasyOCR reader for the processing loop, configured specifically for English text recognition:
reader = easyocr.Reader(["en"])
result = reader.readtext(path, paragraph=True)
The Reader is initialized with the ["en"] language pack, and readtext() executes with paragraph=True to group words into logical text blocks. This returns a list of detections where each entry contains the polygon vertices and the recognized string.
Converting Polygons to Axis-Aligned Bounding Boxes
EasyOCR returns text regions as arbitrary quadrilaterals, but Doc2Graph requires axis-aligned bounding boxes for graph node positioning. The code extracts the first and third polygon points to define the rectangle:
box = [
int(r[0][0][0]), int(r[0][0][1]), # Top-left (x0, y0)
int(r[0][2][0]), int(r[0][2][1]) # Bottom-right (x1, y1)
]
boxes.append(box)
texts.append(r[1])
This conversion generates standard [x0, y0, x1, y1] coordinates stored in the boxes list, while the corresponding recognized strings populate the texts list.
Populating Node Features
After processing all detections, the method aggregates the OCR results into feature dictionaries that the graph construction engine consumes:
features["boxs"] = boxes
features["texts"] = texts
These arrays become the node attributes for the resulting graph, where each text region represents a distinct node with spatial and textual features.
Graph Construction from OCR Data
With features["boxs"] and features["texts"] populated, the GraphBuilder proceeds to wire nodes into edges based on the configured edge_type parameter (defined in doc2graph/data/preprocessing.py). The library supports two connectivity modes:
fully: Creates fully-connected graphs where every text node connects to every other nodeknn: Connects each node to its k-nearest spatial neighbors
The resulting dgl.graph objects encapsulate the document structure, with EasyOCR-derived bounding boxes serving as node coordinates and text strings as node labels. These graphs feed directly into the GNN model defined in doc2graph/models/graphs.py for relationship prediction.
Running Custom Inference: Complete Example
To execute custom inference on a batch of images, initialize the pipeline with pretrained weights and image paths:
from doc2graph.inference import inference
# Load pretrained checkpoint (must exist in CHECKPOINTS directory)
weights = ["funsd-bert.pt"]
# Define image paths for processing
image_paths = [
"samples/invoice1.png",
"samples/invoice2.jpg",
]
# Execute inference (device=-1 for CPU, or specify GPU ID)
inference(weights, image_paths, device=-1)
Behind the scenes, this invocation triggers the EasyOCR extraction flow in __fromIMG(), builds the graph representations, and returns predictions visualized on the original images.
Summary
- EasyOCR is the exclusive text extraction engine for custom image inference in Doc2Graph, invoked inside
GraphBuilder.__fromIMG()indoc2graph/data/graph_builder.py. - Bounding box conversion transforms EasyOCR's polygon outputs into axis-aligned
[x0, y0, x1, y1]coordinates stored infeatures["boxs"]. - Node feature population maps recognized text strings to
features["texts"], establishing the textual attributes for each graph node. - Graph wiring connects OCR-derived nodes using
fullyorknnedge strategies before feeding the structure to the GNN model.
Frequently Asked Questions
What OCR engine does Doc2Graph use for custom images?
Doc2Graph uses EasyOCR as its sole text extraction engine when processing custom images through the inference() function. The implementation specifically calls easyocr.Reader(["en"]) inside doc2graph/data/graph_builder.py to detect and recognize text regions before converting them into graph nodes.
How does Doc2Graph convert EasyOCR output into graph nodes?
The library converts EasyOCR's polygon-based detections into axis-aligned bounding boxes by extracting the first and third vertices of each quadrilateral to form [x0, y0, x1, y1] coordinates. These boxes populate features["boxs"] while the recognized text strings populate features["texts"], creating the node attributes for the document graph.
Can Doc2Graph extract text in languages other than English?
Currently, the EasyOCR reader is hardcoded to ["en"] (English) in doc2graph/data/graph_builder.py. However, the codebase structure suggests future multilingual support could be implemented by modifying the language list passed to easyocr.Reader() and ensuring corresponding model weights are available.
Where is the text extraction logic located in the Doc2Graph codebase?
The core text extraction logic resides in doc2graph/data/graph_builder.py within the GraphBuilder.__fromIMG() method (lines 20-46). This method instantiates the EasyOCR reader, processes images, converts polygons to bounding boxes, and prepares the feature dictionaries used for graph construction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →