How to Configure Node Granularity (gt, yolo, ocr) for Document Graphs in doc2graph
Set node granularity using the --node-granularity flag with values gt, yolo, or ocr to control whether graph nodes represent ground-truth entities, YOLO detections, or OCR words.
The doc2graph library converts documents into graph representations where node granularity determines what each node represents—annotated fields, detected objects, or raw text. This configuration is controlled through command-line arguments and YAML configuration files, with the core logic implemented in the GraphBuilder class. Understanding these three granularity levels allows you to process everything from annotated datasets like FUNSD to raw image collections without annotations.
Understanding Node Granularity Levels
Node granularity defines the fundamental unit of your document graph. The three supported modes in doc2graph serve different data availability scenarios:
gt(Ground Truth): Uses labeled entities from annotation files (JSON or XML). Each node represents an annotated field with predefined labels. Ideal for FUNSD or PAU datasets where ground-truth files exist.yolo: Uses bounding boxes from a pre-trained YOLO detector. Each node represents a detected object (e.g., key-value pairs, tables). Requires pre-computed YOLO prediction files.ocr: Uses individual words detected by EasyOCR. Each node represents a text word extracted directly from raw images. Works with any image dataset without requiring annotations.
Configuration Methods
You can configure node granularity through command-line arguments or by modifying configuration files directly.
Command Line Interface
The --node-granularity flag is defined in doc2graph/main.py (lines 70-74) and accepts one of the three string values:
python -m doc2graph.main --node-granularity gt --src-data FUNSD
This argument is processed by set_preprocessing() in doc2graph/utils.py (lines 86-98), which writes the value to a runtime configuration file that overrides the base settings.
Configuration Files
Default options are specified in configs/base.yaml (lines 23-26):
GRAPHS:
node_granularity:
- gt
# - ocr
# - yolo
While you can edit this file directly, the recommended workflow uses CLI flags to generate a specific preprocessing.yaml via the utility functions, ensuring your configuration is version-controlled and reproducible.
Implementation Details in GraphBuilder
The GraphBuilder class in doc2graph/data/graph_builder.py implements the granularity selection logic. At initialization (line 24), it retrieves the setting:
self.node_granularity = self.cfg_preprocessing.GRAPHS.node_granularity
The class then branches into three distinct processing pipelines based on this value.
Ground Truth Branch (gt)
When node_granularity == "gt", the builder parses annotation files directly:
- FUNSD: Reads
adjusted_annotations/*.json(lines 39-96) - PAU: Reads
<doc>_gt.xmland<doc>_ocr.xml(lines 68-98)
Nodes are created from the annotated entities with labels derived directly from the ground-truth files.
YOLO Detection Branch (yolo)
When node_granularity == "yolo", the builder loads pre-computed predictions:
load_predictions(lines 71-74) reads bounding boxes fromsrc/yolo_bbox/- Nodes represent YOLO detections with class labels from the YOLO output
- Requires prediction files matching the image naming convention
OCR Branch (ocr)
When node_granularity == "ocr" (the else branch), the builder processes raw images:
- Instantiates
easyocr.Readerfor each image - Creates one node per word detected by EasyOCR
- Works with image-only datasets requiring no pre-existing annotations
All three branches converge to the same downstream pipeline for edge generation and feature extraction, differing only in how node_labels and bounding boxes (boxs) are produced.
Practical Usage by Document Type
Select granularity based on your dataset's available assets:
| Document Type | Recommended Granularity | Required Inputs |
|---|---|---|
| FUNSD (forms with JSON annotations) | gt |
images/ + adjusted_annotations/*.json |
| PAU (scanned documents with XML) | gt |
*.tif + <doc>_gt.xml + <doc>_ocr.xml |
| Custom image datasets (no annotations) | ocr |
Raw images only |
| YOLO-processed documents | yolo |
images/ + yolo_bbox/ with prediction files |
Code Examples
Running with Ground Truth Nodes
Process FUNSD using the default ground-truth annotations:
python -m doc2graph.main \
--src-data FUNSD \
--node-granularity gt \
--edge-type fully \
--model e2e \
--gpu 0
Running with OCR Words
Process custom images without annotations:
python -m doc2graph.main \
--src-data CUSTOM \
--data-type img \
--node-granularity ocr \
--edge-type knn \
--model e2e \
--gpu 0
Running with YOLO Detections
Use pre-computed YOLO bounding boxes:
python -m doc2graph.main \
--src-data CUSTOM \
--data-type img \
--node-granularity yolo \
--edge-type fully \
--model e2e \
--gpu 0
Programmatic Configuration
Generate configuration files directly in Python:
from doc2graph.utils import set_preprocessing
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("--node-granularity", type=str, default="yolo")
args = parser.parse_args()
set_preprocessing(args) # Creates preprocessing.yaml with specified granularity
Summary
- Three granularity levels:
gt(annotations),yolo(detections), andocr(text words) control what constitutes a graph node. - CLI configuration: Use
--node-granularityparsed indoc2graph/main.pyand stored viadoc2graph/utils.py. - Implementation:
GraphBuilderindoc2graph/data/graph_builder.pybranches into three pipelines based on the configuration value. - File locations: Defaults live in
configs/base.yaml; runtime values are stored in generated preprocessing YAML files. - Selection criteria: Use
gtfor annotated datasets (FUNSD, PAU),ocrfor raw images, andyolowhen you have pre-computed object detections.
Frequently Asked Questions
What is the default node granularity if I do not specify the flag?
The default value is gt (ground truth), as defined in configs/base.yaml lines 23-26. If your dataset lacks annotation files and you run with the default, the pipeline will fail when attempting to locate the expected JSON or XML ground-truth files.
Can I switch between granularity modes without changing the codebase?
Yes. The granularity is strictly a runtime configuration. Pass --node-granularity ocr or --node-granularity yolo to doc2graph/main.py without modifying any source files. The set_preprocessing utility dynamically generates the configuration file required by the GraphBuilder.
Where does the OCR branch get its text data?
The OCR branch instantiates easyocr.Reader within doc2graph/data/graph_builder.py and processes each image individually. It does not require external annotation files; all text and bounding box data comes directly from the EasyOCR inference results at preprocessing time.
Is it possible to use YOLO granularity without pre-computed bounding boxes?
No. The yolo granularity mode expects existing prediction files in the yolo_bbox/ directory. The load_predictions method (lines 71-74 in graph_builder.py) reads these files to create nodes. You must run your YOLO model separately to generate these detections before invoking doc2graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →