How Doc2Graph Handles Layout Analysis and Table Detection: A Technical Deep Dive
Doc2Graph converts scanned documents into graph structures where geometric edge features capture spatial layout relationships, and a binary edge classifier detects tables by identifying connected components of "table" edges between OCR tokens.
Doc2Graph is an open-source framework that transforms document images and PDFs into graph representations suitable for Graph Neural Networks (GNNs). Understanding how Doc2Graph performs layout analysis and table detection requires examining its graph construction pipeline in GraphBuilder and feature enrichment in FeatureBuilder. The system combines OCR tokenization with geometric feature engineering to jointly learn document structure and tabular region identification according to the andreagemelli/doc2graph source code.
Layout Analysis Through Graph Topology and Geometric Features
Token Extraction and Graph Construction
Layout analysis begins when the OCR engine (easyocr) returns a list of bounding boxes (boxs) and associated texts (texts) for each input image. The GraphBuilder class constructs the initial topology using either a fully-connected strategy via fully_connected or a k-NN neighborhood approach via knn_connection. Both functions return source (u) and destination (v) node indices that define the graph edges.
In doc2graph/data/graph_builder.py, the __fromIMG method handles this edge creation in lines 48‑54.
Spatial Feature Engineering
Once the graph topology exists, FeatureBuilder.add_features enriches nodes and edges with geometric information. Each node receives normalized geometry features (geom) representing the four coordinates of its bounding box divided by image width and height (self.sg).
For edges, the system computes Euclidean distance and angle between connected nodes using the polar function. These values are normalized and quantized into polar bins via to_bin, then stored as edge features in g.edata["feat"]. This allows the GNN to reason about directional relationships such as "this token is to the right of that token."
These operations appear in doc2graph/data/feature_builder.py at lines 71‑75 for node geometry and lines 33‑55 for edge weights.
Table Detection via Edge Classification
Ground Truth Labeling from XML Annotations
Table detection relies on edge-level supervision derived from dataset annotations. In the PAU and FUNSD datasets, XML entries contain <TextRegion> tags with labels "positions" or "total" that identify table cells. During graph construction in __fromPAU, each OCR token is assigned a node label (nl) based on which region it falls into geometrically (if r[0] < c[0] < r[2] …).
This mapping logic resides in doc2graph/data/graph_builder.py at lines 80‑88.
Edge Label Generation
After node labeling, the system generates binary edge targets. If both endpoint nodes belong to a table region (nl[e[0]] == "positions" or nl[e[0]] == "total"), the edge label is set to "table"; otherwise, it is "none". These labels become the edge-class target during model training, where edge_num_classes = 2 in the pretraining configuration.
This labeling mechanism appears in the same file at lines 96‑114.
Inference and Bounding Box Extraction
During inference, the model predicts edge classes, and table regions are extracted by analyzing edges classified as class 1 (table). The inference pipeline:
- Collects table edges:
links = (epreds == 1).nonzero(as_tuple=True)[0] - Builds an induced subgraph:
dgl.edge_subgraph - Computes bounding boxes from the min/max of node geometries within the subgraph
- Scales coordinates back to pixel space
This extraction logic is implemented in doc2graph/training/pau.py at lines 66‑84.
Evaluation Metrics
Predicted table bounding boxes are compared against ground-truth regions using the match_pred_w_gt helper function, which implements IoU-based matching. The system accumulates precision, recall, and F1 scores per-sample and aggregates them across the dataset (all_precisions, all_recalls, all_f1).
The metric computation appears in doc2graph/training/pau.py at lines 49‑63.
End-to-End Implementation Example
The following example demonstrates the complete pipeline from image input to graph-based inference:
from doc2graph.data.graph_builder import GraphBuilder
from doc2graph.data.feature_builder import FeatureBuilder
from doc2graph.inference import inference
# 1️⃣ Build graphs from a list of image/PDF paths
gb = GraphBuilder()
graphs, _, _, feats = gb.get_graph(["sample1.png", "sample2.pdf"], "CUSTOM")
# 2️⃣ Enrich graphs with layout and visual features
fb = FeatureBuilder(d="cpu")
chunks, _ = fb.add_features(graphs, feats)
# 3️⃣ Run inference (weights must be a checkpoint list)
# e.g. ["funsd-e2e-20230101-1234.pt"]
preds = inference(weights=["funsd-e2e-20230101-1234.pt"],
paths=["sample1.png", "sample2.pdf"],
device=0) # GPU 0
The inference function visualizes predicted key-value pairs, while the edge class "table" remains stored in the graph metadata for downstream extraction.
Summary
- Layout analysis in Doc2Graph relies on geometric node features (normalized bounding boxes) and polar edge features (distance and angle) computed in
FeatureBuilder. - Graph construction supports both fully-connected and k-NN topologies via
GraphBuilder.__fromIMG. - Table detection is formulated as binary edge classification, where edges connecting table cells are labeled "table" and others "none."
- Inference extracts tables by finding connected components of predicted table edges and computing their bounding boxes from node geometries.
- Evaluation uses IoU-based matching in
training/pau.pyto compute precision, recall, and F1 for table detection tasks.
Frequently Asked Questions
How does Doc2Graph represent spatial relationships between text tokens?
Doc2Graph represents spatial relationships through polar edge features computed in FeatureBuilder.add_features. The system calculates the Euclidean distance and angle between connected nodes, normalizes these values, and quantizes them into bins. These features allow the GNN to learn directional and proximity-based relationships, such as whether one token appears above or to the left of another.
What distinguishes table edges from non-table edges in the graph?
Table edges are distinguished by node membership in annotated table regions. During training data preparation in GraphBuilder.__fromPAU, nodes receive labels based on whether they fall within XML-annotated table regions (labeled "positions" or "total"). An edge is labeled "table" only if both endpoints reside within such regions; otherwise, it receives the "none" label. The model learns to predict this binary classification during training.
How does the model convert edge predictions into table bounding boxes?
After inference, the model identifies all edges predicted as class 1 (table) and constructs an induced subgraph using dgl.edge_subgraph. The bounding box for each predicted table is derived by computing the minimum and maximum coordinates (min and max) of the node geometries within this subgraph. These normalized coordinates are then scaled back to the original image dimensions to produce the final pixel-level bounding box.
Which datasets support the table detection training pipeline?
The table detection training pipeline primarily supports the PAU and FUNSD datasets. These datasets provide XML annotations containing <TextRegion> entries specifically labeled for table content. The GraphBuilder class parses these annotations in __fromPAU to generate ground-truth node and edge labels necessary for supervised learning of table boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →