How Doc2Graph Saves Inference Results as JSON and Visualization Images

Doc2Graph saves inference results by extracting predicted key-value pairs into a JSON file and overlaying relationship links on the original document as a PNG image, storing both artifacts in the repository's inference/ folder.

The andreagemelli/doc2graph repository provides an end-to-end pipeline for document understanding that automatically persists model predictions. When processing documents, the system generates both machine-readable structured data and human-readable visualizations to validate extraction accuracy. Understanding how Doc2Graph saves inference results as JSON and visualization images is essential for integrating the tool into production document processing workflows.

The Inference Saving Pipeline

The saving mechanism triggers automatically during the inference() routine in doc2graph/inference.py. The workflow processes each input document through four distinct phases before persisting artifacts at lines 95-97.

Graph Construction and Feature Enrichment

First, the GraphBuilder class in doc2graph/data/graph_builder.py (called at lines 26-27) constructs a DGL graph from the input document. Subsequently, the FeatureBuilder in doc2graph/data/feature_builder.py (lines 31-32) enriches nodes and edges with geometric, visual, textual-embedding, and histogram features based on CLI flags.

The system loads the specified checkpoint via model.load_state_dict at line 43. During forward passes (lines 54-56), the model predicts node-class scores (n) and edge-class scores (e). The code extracts the highest-scoring class for each edge as epreds, then filters for relationships where the predicted class equals 1 (the "link" class) at line 59:

links = (epreds == 1).nonzero(as_tuple=True)[0].tolist()

JSON Serialization

For every predicted link, the system constructs dictionaries containing the key and value text along with their bounding boxes (lines 68-73). These dictionaries accumulate in a result list that represents the logical extraction structure [{ "key": {...}, "value": {...} }, …].

Image Visualization

Using Pillow, the code draws relationship visualizations (lines 75-93). A line connects the center points of the two bounding boxes, with the key endpoint rendered in green and the value endpoint in red for immediate visual verification.

File Persistence

Finally, the system writes the artifacts to disk (lines 95-97). The annotated image saves as <name>.png and the JSON list as <name>.json inside the INFERENCE folder defined in doc2graph/paths.py as INFERENCE = ROOT / "inference".

Running Inference from the Command Line

Execute the pipeline using the module's CLI interface:

python -m doc2graph.main \
    --inference \
    --weights e2e-funsd-best.pt \
    --docs path/to/doc1.png path/to/doc2.png \
    --gpu 0

This command triggers the complete workflow. Upon completion, the inference/ directory contains paired outputs for each document:


inference/
├── doc1.png   # visualization with coloured links

├── doc1.json  # extracted key-value pairs

├── doc2.png
└── doc2.json

Programmatic Inference in Python

Import the inference function directly to process documents within Python applications:

from doc2graph.inference import inference

# List of absolute file paths to the documents you want to process

doc_paths = [
    "/absolute/path/to/form1.png",
    "/absolute/path/to/form2.png",
]

# Name of the checkpoint file placed under doc2graph/models/checkpoints/

weights = ["e2e-funsd-best.pt"]

# Run on CPU (device = -1) or specify a GPU id

inference(weights, doc_paths, device=-1)

After execution, the function creates the same JSON and PNG files in the inference/ directory without requiring shell interaction.

Output File Formats

JSON Structure

The generated JSON files contain lists of key-value pair dictionaries. Each entry specifies the text content and spatial coordinates:

import json
from pathlib import Path

result_path = Path("inference/doc1.json")
with result_path.open() as f:
    data = json.load(f)

print(data[0])

# Output example:

# {

#   "key": {"text": "Invoice No.", "box": [100, 150, 200, 180]},

#   "value": {"text": "12345", "box": [210, 150, 260, 180]}

# }

PNG Visualization Details

The PNG images provide human-readable verification by overlaying colored connection lines on the original document. Green endpoints mark key entities while red endpoints indicate corresponding values, making it easy to visually validate the model's relationship predictions.

Summary

  • Doc2Graph automatically generates both JSON extractions and PNG visualizations for every processed document during inference.
  • The inference() function in doc2graph/inference.py orchestrates the pipeline, from graph construction through file persistence at lines 95-97.
  • JSON files store structured key-value pairs with bounding box coordinates in the inference/ directory.
  • PNG images use color-coded overlays (green for keys, red for values) to visualize predicted relationships.
  • Output paths derive from the INFERENCE constant defined in doc2graph/paths.py.

Frequently Asked Questions

Where does Doc2Graph save inference results by default?

Doc2Graph saves all inference artifacts in the inference/ folder relative to the repository root. This path is defined as INFERENCE = ROOT / "inference" in doc2graph/paths.py, and both JSON and PNG files for each processed document appear in this directory with matching base filenames.

What information does the inference JSON file contain?

The JSON file contains a list of dictionaries, where each dictionary represents a predicted key-value relationship. Every entry includes the key text and bounding box coordinates, paired with the corresponding value text and its bounding box, enabling precise spatial and semantic extraction from the source document.

How does the visualization image represent predicted relationships?

The PNG visualization uses Pillow to draw lines connecting the center points of predicted key-value pairs. The key endpoint renders in green while the value endpoint renders in red, providing immediate visual feedback on the model's link predictions overlaid on the original document image.

Can I customize the output directory for inference results?

The current implementation in doc2graph/paths.py hardcodes the INFERENCE path as ROOT / "inference". To change the output location, you must modify the INFERENCE constant in doc2graph/paths.py or manually move files after generation, as the CLI and programmatic interfaces do not expose a configurable output path parameter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →