What Features Can Be Extracted and Added to Graph Nodes in doc2graph: geom, embs, hist, and visual

The doc2graph library extracts four primary feature types—geometric coordinates (geom), spaCy text embeddings (embs), character-type histograms (hist), and UNet visual embeddings (visual)—and concatenates them into DGL graph node attributes via the FeatureBuilder class.

In document intelligence pipelines, the quality of graph representations depends heavily on the richness of node features. The doc2graph repository implements a modular feature extraction system where each detected text box becomes a graph node annotated with multi-modal data, all controlled through boolean flags in preprocessing.yaml.

The Four Primary Node Feature Categories

The FeatureBuilder class in doc2graph/data/feature_builder.py constructs node feature vectors by conditionally appending components based on configuration flags. Each flag corresponds to a specific mathematical representation of the document content.

Geometric Features (geom)

When add_geom is enabled, the system extracts normalized bounding-box coordinates for every detected text region. The x, y, w, h values are normalized by image width and height (self.sg) to ensure scale invariance across different document resolutions. These four-dimensional vectors are stored separately in g.ndata["geom"] and concatenated into the main feature tensor.

According to the source code in doc2graph/data/feature_builder.py (lines 78–85), the normalization occurs during the add_features method execution, ensuring all spatial coordinates fall within a consistent [0, 1] range regardless of original image dimensions.

Text Embeddings (embs)

The add_embs flag triggers semantic encoding using spaCy's large English model (en_core_web_lg). Each tokenized text string generates a dense 300-dimensional vector that captures semantic meaning beyond raw character sequences. The initialization occurs in FeatureBuilder.__init__ (lines 34–38), while the embedding loop processes each node's text content in lines 95–103.

These embeddings provide contextual understanding of entity types (e.g., distinguishing between company names, dates, and monetary values) without requiring task-specific fine-tuning.

Character Histograms (hist)

Enabling add_hist appends a 4-dimensional character-type distribution to each node. The helper function get_histogram in doc2graph/data/utils.py (lines 80–92) categorizes every character in the token into four bins:

  • Letters (alphabetic characters)
  • Digits (numeric characters)
  • Symbols (punctuation and special characters)
  • Others (whitespace and unclassified characters)

Each bin contains the normalized percentage of that character type, creating a compact linguistic signature that helps distinguish between purely numeric fields, alphabetic headers, and mixed-content paragraphs.

Visual Embeddings (visual)

The add_visual flag activates the most computationally intensive feature extraction pathway. A pretrained UNet encoder (utilizing a mobilenet_v2 backbone defined in doc2graph/models/unet/model.py) processes the entire document image. For each bounding box, ROI-Align extracts a fixed-size feature map from the encoder's output, which is then flattened into a dense vector.

As implemented in doc2graph/data/feature_builder.py (lines 38–50 for encoder initialization, lines 106–131 for ROI-Align extraction), this captures visual patterns such as font styles, background textures, and layout context that pure text analysis would miss.

How FeatureBuilder Orchestrates Extraction

The extraction pipeline follows a strict configuration-driven architecture. During initialization, FeatureBuilder loads boolean flags from preprocessing.yaml via utils.get_config:

self.add_geom   = self.cfg_preprocessing.FEATURES.add_geom
self.add_embs   = self.cfg_preprocessing.FEATURES.add_embs
self.add_hist   = self.cfg_preprocessing.FEATURES.add_hist
self.add_visual = self.cfg_preprocessing.FEATURES.add_visual

Within the add_features method, the system iterates through all graphs and concatenates enabled features into a single tensor stored at g.ndata["feat"]. The dimensionality of this tensor varies based on the active flags: 4 for geometry, 300 for text embeddings, 4 for histograms, and variable dimensions (depending on UNet decoder depth) for visual features.

Edge Weight Features (add_eweights)

While the primary question focuses on node attributes, the related add_eweights flag computes polar geometric relationships between node pairs. The polar function in doc2graph/data/utils.py (lines 8–58) calculates Euclidean distance and relative angle between bounding boxes. The to_bin function (lines 31–67) discretizes these continuous values into binary-encoded bins, storing the results as edge attributes in g.edata["feat"] for relational reasoning.

Practical Implementation Example

The following workflow demonstrates how to initialize the feature extraction pipeline and inspect the resulting node tensors:

from doc2graph.data.feature_builder import FeatureBuilder
from doc2graph.data.graph_builder import GraphBuilder
import torch

# Build initial graphs from FUNSD dataset

gb = GraphBuilder()
graphs, node_labels, edge_labels, feats = gb.get_graph(
    src_path="/path/to/FUNSD", src_data="FUNSD"
)

# Initialize feature extractor with CPU or CUDA device

fb = FeatureBuilder(d="cpu")

# Add configured features to each graph

feats_per_node, feat_dim = fb.add_features(graphs, feats)

# Inspect resulting node attributes

g = graphs[0]
print("Geometric shape:", g.ndata["geom"].shape)      # (num_nodes, 4)

print("Concatenated features:", g.ndata["feat"].shape) # (num_nodes, total_dim)

print("Edge attributes:", g.edata["feat"].shape)       # (num_edges, 2*b) if eweights enabled

This example shows the complete pipeline: GraphBuilder creates the initial DGL structure from document sources, then FeatureBuilder enriches each node with the multi-modal features specified in the preprocessing configuration.

Summary

  • Geometric (geom) features provide 4-dimensional normalized bounding-box coordinates (x, y, w, h) for spatial layout understanding.
  • Text embeddings (embs) deliver 300-dimensional semantic vectors from spaCy en_core_web_lg for linguistic content analysis.
  • Histograms (hist) supply 4-dimensional character-type distributions distinguishing letters, digits, symbols, and other characters.
  • Visual embeddings (visual) generate dense convolutional features via UNet encoder and ROI-Align for appearance-based reasoning.
  • All features are concatenated into g.ndata["feat"] by FeatureBuilder.add_features() in doc2graph/data/feature_builder.py, with configuration controlled through preprocessing.yaml flags.

Frequently Asked Questions

How do I enable specific features in doc2graph?

Modify the preprocessing.yaml configuration file to set the boolean flags under the FEATURES section. Set add_geom: true, add_embs: true, add_hist: true, or add_visual: true depending on which modalities your graph neural network requires. The FeatureBuilder class automatically detects these settings during initialization via utils.get_config.

What is the total dimensionality of node features when all flags are enabled?

The final dimensionality equals the sum of all active feature components: 4 (geometry) + 300 (spaCy embeddings) + 4 (histogram) + N (visual embedding size, determined by UNet architecture). The visual embedding dimension depends on the specific UNet decoder configuration in doc2graph/models/unet/model.py, typically producing several hundred additional dimensions.

Does doc2graph support custom visual encoders beyond the default UNet?

The current implementation in doc2graph/data/feature_builder.py (lines 38–50) instantiates a specific UNet encoder with MobileNetV2 backbone for visual feature extraction. While the source code architecture supports replacing this encoder, doing so requires modifying the FeatureBuilder initialization logic and ensuring compatible tensor dimensions for the ROI-Align operation in lines 106–131.

How are geometric features normalized in the FeatureBuilder?

During the add_features method execution (lines 78–85 of doc2graph/data/feature_builder.py), raw bounding-box coordinates are divided by image width and height (self.sg). This normalization ensures that x and w values are scaled relative to image width, while y and h values scale relative to image height, producing consistent [0, 1] range coordinates regardless of original document resolution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →