# How to Configure Node Granularity (gt, yolo, ocr) for Document Graphs in doc2graph

> Learn to configure node granularity gt yolo or ocr in doc2graph for document graphs. Control graph nodes based on ground truth YOLO or OCR data.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**Set node granularity using the `--node-granularity` flag with values `gt`, `yolo`, or `ocr` to control whether graph nodes represent ground-truth entities, YOLO detections, or OCR words.**

The **doc2graph** library converts documents into graph representations where **node granularity** determines what each node represents—annotated fields, detected objects, or raw text. This configuration is controlled through command-line arguments and YAML configuration files, with the core logic implemented in the `GraphBuilder` class. Understanding these three granularity levels allows you to process everything from annotated datasets like FUNSD to raw image collections without annotations.

## Understanding Node Granularity Levels

Node granularity defines the fundamental unit of your document graph. The three supported modes in doc2graph serve different data availability scenarios:

- **`gt` (Ground Truth)**: Uses labeled entities from annotation files (JSON or XML). Each node represents an annotated field with predefined labels. Ideal for FUNSD or PAU datasets where ground-truth files exist.
- **`yolo`**: Uses bounding boxes from a pre-trained YOLO detector. Each node represents a detected object (e.g., key-value pairs, tables). Requires pre-computed YOLO prediction files.
- **`ocr`**: Uses individual words detected by EasyOCR. Each node represents a text word extracted directly from raw images. Works with any image dataset without requiring annotations.

## Configuration Methods

You can configure node granularity through command-line arguments or by modifying configuration files directly.

### Command Line Interface

The `--node-granularity` flag is defined in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) (lines 70-74) and accepts one of the three string values:

```bash
python -m doc2graph.main --node-granularity gt --src-data FUNSD

```

This argument is processed by `set_preprocessing()` in [`doc2graph/utils.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/utils.py) (lines 86-98), which writes the value to a runtime configuration file that overrides the base settings.

### Configuration Files

Default options are specified in [`configs/base.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/configs/base.yaml) (lines 23-26):

```yaml
GRAPHS:
  node_granularity:
    - gt
    # - ocr

    # - yolo

```

While you can edit this file directly, the recommended workflow uses CLI flags to generate a specific [`preprocessing.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/preprocessing.yaml) via the utility functions, ensuring your configuration is version-controlled and reproducible.

## Implementation Details in GraphBuilder

The `GraphBuilder` class in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) implements the granularity selection logic. At initialization (line 24), it retrieves the setting:

```python
self.node_granularity = self.cfg_preprocessing.GRAPHS.node_granularity

```

The class then branches into three distinct processing pipelines based on this value.

### Ground Truth Branch (gt)

When `node_granularity == "gt"`, the builder parses annotation files directly:

- **FUNSD**: Reads `adjusted_annotations/*.json` (lines 39-96)
- **PAU**: Reads `<doc>_gt.xml` and `<doc>_ocr.xml` (lines 68-98)

Nodes are created from the annotated entities with labels derived directly from the ground-truth files.

### YOLO Detection Branch (yolo)

When `node_granularity == "yolo"`, the builder loads pre-computed predictions:

- `load_predictions` (lines 71-74) reads bounding boxes from `src/yolo_bbox/`
- Nodes represent YOLO detections with class labels from the YOLO output
- Requires prediction files matching the image naming convention

### OCR Branch (ocr)

When `node_granularity == "ocr"` (the else branch), the builder processes raw images:

- Instantiates `easyocr.Reader` for each image
- Creates one node per word detected by EasyOCR
- Works with image-only datasets requiring no pre-existing annotations

All three branches converge to the same downstream pipeline for edge generation and feature extraction, differing only in how `node_labels` and bounding boxes (`boxs`) are produced.

## Practical Usage by Document Type

Select granularity based on your dataset's available assets:

| Document Type | Recommended Granularity | Required Inputs |
|--------------|------------------------|----------------|
| **FUNSD** (forms with JSON annotations) | `gt` | `images/` + `adjusted_annotations/*.json` |
| **PAU** (scanned documents with XML) | `gt` | `*.tif` + `<doc>_gt.xml` + `<doc>_ocr.xml` |
| **Custom image datasets** (no annotations) | `ocr` | Raw images only |
| **YOLO-processed documents** | `yolo` | `images/` + `yolo_bbox/` with prediction files |

## Code Examples

### Running with Ground Truth Nodes

Process FUNSD using the default ground-truth annotations:

```bash
python -m doc2graph.main \
    --src-data FUNSD \
    --node-granularity gt \
    --edge-type fully \
    --model e2e \
    --gpu 0

```

### Running with OCR Words

Process custom images without annotations:

```bash
python -m doc2graph.main \
    --src-data CUSTOM \
    --data-type img \
    --node-granularity ocr \
    --edge-type knn \
    --model e2e \
    --gpu 0

```

### Running with YOLO Detections

Use pre-computed YOLO bounding boxes:

```bash
python -m doc2graph.main \
    --src-data CUSTOM \
    --data-type img \
    --node-granularity yolo \
    --edge-type fully \
    --model e2e \
    --gpu 0

```

### Programmatic Configuration

Generate configuration files directly in Python:

```python
from doc2graph.utils import set_preprocessing
import argparse

parser = argparse.ArgumentParser()
parser.add_argument("--node-granularity", type=str, default="yolo")
args = parser.parse_args()

set_preprocessing(args)  # Creates preprocessing.yaml with specified granularity

```

## Summary

- **Three granularity levels**: `gt` (annotations), `yolo` (detections), and `ocr` (text words) control what constitutes a graph node.
- **CLI configuration**: Use `--node-granularity` parsed in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) and stored via [`doc2graph/utils.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/utils.py).
- **Implementation**: `GraphBuilder` in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) branches into three pipelines based on the configuration value.
- **File locations**: Defaults live in [`configs/base.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/configs/base.yaml); runtime values are stored in generated preprocessing YAML files.
- **Selection criteria**: Use `gt` for annotated datasets (FUNSD, PAU), `ocr` for raw images, and `yolo` when you have pre-computed object detections.

## Frequently Asked Questions

### What is the default node granularity if I do not specify the flag?

The default value is `gt` (ground truth), as defined in [`configs/base.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/configs/base.yaml) lines 23-26. If your dataset lacks annotation files and you run with the default, the pipeline will fail when attempting to locate the expected JSON or XML ground-truth files.

### Can I switch between granularity modes without changing the codebase?

Yes. The granularity is strictly a runtime configuration. Pass `--node-granularity ocr` or `--node-granularity yolo` to [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) without modifying any source files. The `set_preprocessing` utility dynamically generates the configuration file required by the `GraphBuilder`.

### Where does the OCR branch get its text data?

The OCR branch instantiates `easyocr.Reader` within [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) and processes each image individually. It does not require external annotation files; all text and bounding box data comes directly from the EasyOCR inference results at preprocessing time.

### Is it possible to use YOLO granularity without pre-computed bounding boxes?

No. The `yolo` granularity mode expects existing prediction files in the `yolo_bbox/` directory. The `load_predictions` method (lines 71-74 in [`graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/graph_builder.py)) reads these files to create nodes. You must run your YOLO model separately to generate these detections before invoking doc2graph.