How to Configure Node Granularity (gt, yolo, ocr) for Document Graphs in doc2graph

Set node granularity using the --node-granularity flag with values gt, yolo, or ocr to control whether graph nodes represent ground-truth entities, YOLO detections, or OCR words.

The doc2graph library converts documents into graph representations where node granularity determines what each node represents—annotated fields, detected objects, or raw text. This configuration is controlled through command-line arguments and YAML configuration files, with the core logic implemented in the GraphBuilder class. Understanding these three granularity levels allows you to process everything from annotated datasets like FUNSD to raw image collections without annotations.

Understanding Node Granularity Levels

Node granularity defines the fundamental unit of your document graph. The three supported modes in doc2graph serve different data availability scenarios:

  • gt (Ground Truth): Uses labeled entities from annotation files (JSON or XML). Each node represents an annotated field with predefined labels. Ideal for FUNSD or PAU datasets where ground-truth files exist.
  • yolo: Uses bounding boxes from a pre-trained YOLO detector. Each node represents a detected object (e.g., key-value pairs, tables). Requires pre-computed YOLO prediction files.
  • ocr: Uses individual words detected by EasyOCR. Each node represents a text word extracted directly from raw images. Works with any image dataset without requiring annotations.

Configuration Methods

You can configure node granularity through command-line arguments or by modifying configuration files directly.

Command Line Interface

The --node-granularity flag is defined in doc2graph/main.py (lines 70-74) and accepts one of the three string values:

python -m doc2graph.main --node-granularity gt --src-data FUNSD

This argument is processed by set_preprocessing() in doc2graph/utils.py (lines 86-98), which writes the value to a runtime configuration file that overrides the base settings.

Configuration Files

Default options are specified in configs/base.yaml (lines 23-26):

GRAPHS:
  node_granularity:
    - gt
    # - ocr

    # - yolo

While you can edit this file directly, the recommended workflow uses CLI flags to generate a specific preprocessing.yaml via the utility functions, ensuring your configuration is version-controlled and reproducible.

Implementation Details in GraphBuilder

The GraphBuilder class in doc2graph/data/graph_builder.py implements the granularity selection logic. At initialization (line 24), it retrieves the setting:

self.node_granularity = self.cfg_preprocessing.GRAPHS.node_granularity

The class then branches into three distinct processing pipelines based on this value.

Ground Truth Branch (gt)

When node_granularity == "gt", the builder parses annotation files directly:

  • FUNSD: Reads adjusted_annotations/*.json (lines 39-96)
  • PAU: Reads <doc>_gt.xml and <doc>_ocr.xml (lines 68-98)

Nodes are created from the annotated entities with labels derived directly from the ground-truth files.

YOLO Detection Branch (yolo)

When node_granularity == "yolo", the builder loads pre-computed predictions:

  • load_predictions (lines 71-74) reads bounding boxes from src/yolo_bbox/
  • Nodes represent YOLO detections with class labels from the YOLO output
  • Requires prediction files matching the image naming convention

OCR Branch (ocr)

When node_granularity == "ocr" (the else branch), the builder processes raw images:

  • Instantiates easyocr.Reader for each image
  • Creates one node per word detected by EasyOCR
  • Works with image-only datasets requiring no pre-existing annotations

All three branches converge to the same downstream pipeline for edge generation and feature extraction, differing only in how node_labels and bounding boxes (boxs) are produced.

Practical Usage by Document Type

Select granularity based on your dataset's available assets:

Document Type Recommended Granularity Required Inputs
FUNSD (forms with JSON annotations) gt images/ + adjusted_annotations/*.json
PAU (scanned documents with XML) gt *.tif + <doc>_gt.xml + <doc>_ocr.xml
Custom image datasets (no annotations) ocr Raw images only
YOLO-processed documents yolo images/ + yolo_bbox/ with prediction files

Code Examples

Running with Ground Truth Nodes

Process FUNSD using the default ground-truth annotations:

python -m doc2graph.main \
    --src-data FUNSD \
    --node-granularity gt \
    --edge-type fully \
    --model e2e \
    --gpu 0

Running with OCR Words

Process custom images without annotations:

python -m doc2graph.main \
    --src-data CUSTOM \
    --data-type img \
    --node-granularity ocr \
    --edge-type knn \
    --model e2e \
    --gpu 0

Running with YOLO Detections

Use pre-computed YOLO bounding boxes:

python -m doc2graph.main \
    --src-data CUSTOM \
    --data-type img \
    --node-granularity yolo \
    --edge-type fully \
    --model e2e \
    --gpu 0

Programmatic Configuration

Generate configuration files directly in Python:

from doc2graph.utils import set_preprocessing
import argparse

parser = argparse.ArgumentParser()
parser.add_argument("--node-granularity", type=str, default="yolo")
args = parser.parse_args()

set_preprocessing(args)  # Creates preprocessing.yaml with specified granularity

Summary

  • Three granularity levels: gt (annotations), yolo (detections), and ocr (text words) control what constitutes a graph node.
  • CLI configuration: Use --node-granularity parsed in doc2graph/main.py and stored via doc2graph/utils.py.
  • Implementation: GraphBuilder in doc2graph/data/graph_builder.py branches into three pipelines based on the configuration value.
  • File locations: Defaults live in configs/base.yaml; runtime values are stored in generated preprocessing YAML files.
  • Selection criteria: Use gt for annotated datasets (FUNSD, PAU), ocr for raw images, and yolo when you have pre-computed object detections.

Frequently Asked Questions

What is the default node granularity if I do not specify the flag?

The default value is gt (ground truth), as defined in configs/base.yaml lines 23-26. If your dataset lacks annotation files and you run with the default, the pipeline will fail when attempting to locate the expected JSON or XML ground-truth files.

Can I switch between granularity modes without changing the codebase?

Yes. The granularity is strictly a runtime configuration. Pass --node-granularity ocr or --node-granularity yolo to doc2graph/main.py without modifying any source files. The set_preprocessing utility dynamically generates the configuration file required by the GraphBuilder.

Where does the OCR branch get its text data?

The OCR branch instantiates easyocr.Reader within doc2graph/data/graph_builder.py and processes each image individually. It does not require external annotation files; all text and bounding box data comes directly from the EasyOCR inference results at preprocessing time.

Is it possible to use YOLO granularity without pre-computed bounding boxes?

No. The yolo granularity mode expects existing prediction files in the yolo_bbox/ directory. The load_predictions method (lines 71-74 in graph_builder.py) reads these files to create nodes. You must run your YOLO model separately to generate these detections before invoking doc2graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →