# How the GraphBuilder Class Constructs DGL Graphs from FUNSD and PAU Datasets

> Learn how the GraphBuilder class constructs DGL graphs from FUNSD and PAU datasets by parsing layout data, extracting features, and building edges for advanced document analysis.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**The `GraphBuilder` class in andreagemelli/doc2graph converts FUNSD JSON annotations and PAU XML documents into DGL graph objects by parsing layout data, extracting node features (bounding boxes, text, semantic labels), building either fully-connected or k-NN edge lists, and labeling edges based on dataset-specific semantic relationships before instantiating `dgl.graph` objects.**

The `GraphBuilder` class serves as the core data preparation component in the doc2graph repository, bridging raw document formats and graph neural network inputs. It handles the complete pipeline from parsing FUNSD JSON annotations and PAU XML structures to instantiating `dgl.DGLGraph` objects ready for training. This article examines the implementation details in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) to explain how heterogeneous document layouts become homogeneous graph structures.

## Architecture of the GraphBuilder Pipeline

The class orchestrates the conversion through a standardized workflow defined in `get_graph` (lines 28-55). First, `__init__` loads configuration parameters including `edge_type` and node granularity (lines 19-26). The method then dispatches to private handlers `__fromFUNSD` or `__fromPAU` based on the `src_data` parameter. Both handlers follow a consistent pattern: parse annotations, extract node features, build edge indices, assign edge labels, and construct the DGL graph.

## Processing FUNSD Annotations

### Parsing JSON Structure

For FUNSD documents, the builder reads annotation files from the `adjusted_annotations` directory:

```python
with open(os.path.join(src, "adjusted_annotations", file), "r") as f:
    form = json.load(f)["form"]

```

This extracts the form array containing all document elements with their layout and semantic information.

### Node Feature Extraction

Each element yields four critical attributes: `box` (bounding box coordinates), `text` (OCR content), `label` (semantic class like *header* or *question*), and `id` (unique identifier). The builder also captures `linking` arrays that define ground-truth relationships between nodes, storing these as `pair_labels` for subsequent edge classification.

### Edge Construction and Labeling

The class supports two connectivity schemes. When `self.edge_type == "fully"`, it calls `fully_connected` (lines 100-114) to generate a dense adjacency matrix. For spatial pruning, `knn_connection` selects *k* nearest neighbors based on bounding box proximity. Both methods return parallel index lists `u` and `v`.

The builder then iterates through the edge list to assign labels:

```python
if edge in pair_labels:
    el.append("pair")
else:
    el.append("none")

```

Edges matching the JSON linking annotations receive the **"pair"** label; all others receive **"none"**.

### DGL Graph Instantiation

Finally, the builder creates the graph object in `__fromFUNSD` (lines 88-95):

```python
g = dgl.graph(
    (torch.tensor(u), torch.tensor(v)),
    num_nodes=len(boxs),
    idtype=torch.int32,
)

```

The method returns the graph alongside node labels, edge labels, and feature dictionaries containing `paths`, `texts`, and `boxs`.

## Processing PAU Documents

### XML Parsing and Region Detection

The PAU handler processes Riba dataset documents by pairing `.tif` images with [`_gt.xml`](https://github.com/andreagemelli/doc2graph/blob/main/_gt.xml) (ground truth) and [`_ocr.xml`](https://github.com/andreagemelli/doc2graph/blob/main/_ocr.xml) (OCR) files. It constructs file paths by stripping extensions:

```python
img_name = image.split(".")[0]
file_gt = img_name + "_gt.xml"
file_ocr = img_name + "_ocr.xml"

```

The builder extracts semantic regions from the ground truth XML, storing region labels and bounding boxes in a `regions` list (lines 99-132).

### OCR Token Assignment

The builder parses OCR tokens and assigns each to a region by calculating the word's center point and checking containment against region bounding boxes:

```python
c = center(word_bbox)
for reg in regions:
    if reg[0] <= c[0] <= reg[2] and reg[1] <= c[1] <= reg[3]:
        word_label = reg[0]
        break

```

This produces `tokens_bbox`, `tokens_text`, and `nl` (node labels) arrays used as graph features.

### Edge Labeling Strategy

PAU edges follow the same generation logic as FUNSD (fully-connected or k-NN), but use different semantic labels. Edges connecting nodes within the same **"positions"** or **"total"** regions receive the **"table"** label; all others are **"none"**:

```python
if (nl[e[0]] == nl[e[1]]) and (nl[e[0]] == "positions" or nl[e[0]] == "total"):
    el.append("table")
else:
    el.append("none")

```

### Graph Creation

The DGL graph instantiation in `__fromPAU` (lines 172-182) mirrors the FUNSD approach, using `torch.tensor` conversions for edge indices and setting `num_nodes` to the token count.

## Edge Generation Strategies

The `GraphBuilder` implements two geometric connection schemes in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py). The **fully-connected** approach in `fully_connected` creates complete graphs where every node connects to every other node, preserving all possible relationships but increasing computational cost quadratically. The **k-NN** approach in `knn_connection` constrains connections to spatially proximal neighbors based on projection windows, reducing graph density while maintaining local layout dependencies. Both methods return coordinate lists `u` and `v` compatible with DGL's graph constructor.

## Practical Implementation Examples

### Building FUNSD Graphs

```python
from doc2graph.data.graph_builder import GraphBuilder

builder = GraphBuilder()
graphs, node_labels, edge_labels, features = builder.get_graph(
    src_path="/path/to/FUNSD",
    src_data="FUNSD"
)

g = graphs[0]
print(f"Nodes: {g.num_nodes()}, Edges: {g.num_edges()}")
print(f"Edge label distribution: {set(edge_labels[0])}")

```

### Building PAU Graphs

```python
from doc2graph.data.graph_builder import GraphBuilder

builder = GraphBuilder()
graphs, node_labels, edge_labels, features = builder.get_graph(
    src_path="/path/to/PAU_dataset",
    src_data="PAU"
)

g = graphs[0]
print(f"First 5 node labels: {node_labels[0][:5]}")

```

## Summary

- The `GraphBuilder` class in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) serves as the primary interface for converting FUNSD and PAU documents into DGL graphs.
- For FUNSD, it parses JSON annotations from the `adjusted_annotations` directory, extracting bounding boxes, text, and linking information.
- For PAU, it pairs TIFF images with XML ground truth and OCR files, assigning tokens to semantic regions via center-point containment checks.
- Edge construction supports both **fully-connected** and **k-NN** schemes, with the method determined by the `edge_type` configuration parameter.
- Edge labels are dataset-specific: FUNSD uses "pair"/"none" based on linking annotations, while PAU uses "table"/"none" based on region membership.
- All graphs are instantiated using `dgl.graph()` with `torch.tensor` edge indices and `torch.int32` ID type.

## Frequently Asked Questions

### What file formats does GraphBuilder support?

GraphBuilder processes FUNSD datasets stored as JSON annotations in an `adjusted_annotations` subdirectory and PAU datasets using XML ground truth paired with XML OCR output. For PAU, it expects `.tif` images with corresponding [`_gt.xml`](https://github.com/andreagemelli/doc2graph/blob/main/_gt.xml) and [`_ocr.xml`](https://github.com/andreagemelli/doc2graph/blob/main/_ocr.xml) files following the Riba dataset structure.

### How does the k-NN edge generation differ from fully-connected?

The **k-NN** strategy in `knn_connection` limits edges to the *k* spatially nearest neighbors based on bounding box distances, creating sparse graphs that capture local document layout. **Fully-connected** graphs link every node to every other node, preserving global relationships but significantly increasing edge count and memory consumption.

### Can I customize the edge labeling logic for other datasets?

Yes. The `GraphBuilder` structure allows extension by implementing new private methods similar to `__fromFUNSD` and `__fromPAU`. You would define custom logic to populate relationship arrays (like `pair_labels` or regional constraints), then map these to string labels in the edge list iteration before calling `dgl.graph()`.

### What version of DGL does this implementation require?

The code uses the standard `dgl.graph()` constructor with tuple-formatted edge indices and explicit `idtype=torch.int32`, compatible with DGL 0.6.x and later versions. The repository dependencies typically specify the exact version in [`requirements.txt`](https://github.com/andreagemelli/doc2graph/blob/main/requirements.txt) or [`setup.py`](https://github.com/andreagemelli/doc2graph/blob/main/setup.py).