How Doc2Graph Saves Inference Results as JSON and Visualization Images
Doc2Graph saves inference results by extracting predicted key-value pairs into a JSON file and overlaying relationship links on the original document as a PNG image, storing both artifacts in the repository's inference/ folder.
The andreagemelli/doc2graph repository provides an end-to-end pipeline for document understanding that automatically persists model predictions. When processing documents, the system generates both machine-readable structured data and human-readable visualizations to validate extraction accuracy. Understanding how Doc2Graph saves inference results as JSON and visualization images is essential for integrating the tool into production document processing workflows.
The Inference Saving Pipeline
The saving mechanism triggers automatically during the inference() routine in doc2graph/inference.py. The workflow processes each input document through four distinct phases before persisting artifacts at lines 95-97.
Graph Construction and Feature Enrichment
First, the GraphBuilder class in doc2graph/data/graph_builder.py (called at lines 26-27) constructs a DGL graph from the input document. Subsequently, the FeatureBuilder in doc2graph/data/feature_builder.py (lines 31-32) enriches nodes and edges with geometric, visual, textual-embedding, and histogram features based on CLI flags.
Model Prediction and Link Extraction
The system loads the specified checkpoint via model.load_state_dict at line 43. During forward passes (lines 54-56), the model predicts node-class scores (n) and edge-class scores (e). The code extracts the highest-scoring class for each edge as epreds, then filters for relationships where the predicted class equals 1 (the "link" class) at line 59:
links = (epreds == 1).nonzero(as_tuple=True)[0].tolist()
JSON Serialization
For every predicted link, the system constructs dictionaries containing the key and value text along with their bounding boxes (lines 68-73). These dictionaries accumulate in a result list that represents the logical extraction structure [{ "key": {...}, "value": {...} }, …].
Image Visualization
Using Pillow, the code draws relationship visualizations (lines 75-93). A line connects the center points of the two bounding boxes, with the key endpoint rendered in green and the value endpoint in red for immediate visual verification.
File Persistence
Finally, the system writes the artifacts to disk (lines 95-97). The annotated image saves as <name>.png and the JSON list as <name>.json inside the INFERENCE folder defined in doc2graph/paths.py as INFERENCE = ROOT / "inference".
Running Inference from the Command Line
Execute the pipeline using the module's CLI interface:
python -m doc2graph.main \
--inference \
--weights e2e-funsd-best.pt \
--docs path/to/doc1.png path/to/doc2.png \
--gpu 0
This command triggers the complete workflow. Upon completion, the inference/ directory contains paired outputs for each document:
inference/
├── doc1.png # visualization with coloured links
├── doc1.json # extracted key-value pairs
├── doc2.png
└── doc2.json
Programmatic Inference in Python
Import the inference function directly to process documents within Python applications:
from doc2graph.inference import inference
# List of absolute file paths to the documents you want to process
doc_paths = [
"/absolute/path/to/form1.png",
"/absolute/path/to/form2.png",
]
# Name of the checkpoint file placed under doc2graph/models/checkpoints/
weights = ["e2e-funsd-best.pt"]
# Run on CPU (device = -1) or specify a GPU id
inference(weights, doc_paths, device=-1)
After execution, the function creates the same JSON and PNG files in the inference/ directory without requiring shell interaction.
Output File Formats
JSON Structure
The generated JSON files contain lists of key-value pair dictionaries. Each entry specifies the text content and spatial coordinates:
import json
from pathlib import Path
result_path = Path("inference/doc1.json")
with result_path.open() as f:
data = json.load(f)
print(data[0])
# Output example:
# {
# "key": {"text": "Invoice No.", "box": [100, 150, 200, 180]},
# "value": {"text": "12345", "box": [210, 150, 260, 180]}
# }
PNG Visualization Details
The PNG images provide human-readable verification by overlaying colored connection lines on the original document. Green endpoints mark key entities while red endpoints indicate corresponding values, making it easy to visually validate the model's relationship predictions.
Summary
- Doc2Graph automatically generates both JSON extractions and PNG visualizations for every processed document during inference.
- The
inference()function indoc2graph/inference.pyorchestrates the pipeline, from graph construction through file persistence at lines 95-97. - JSON files store structured key-value pairs with bounding box coordinates in the
inference/directory. - PNG images use color-coded overlays (green for keys, red for values) to visualize predicted relationships.
- Output paths derive from the
INFERENCEconstant defined indoc2graph/paths.py.
Frequently Asked Questions
Where does Doc2Graph save inference results by default?
Doc2Graph saves all inference artifacts in the inference/ folder relative to the repository root. This path is defined as INFERENCE = ROOT / "inference" in doc2graph/paths.py, and both JSON and PNG files for each processed document appear in this directory with matching base filenames.
What information does the inference JSON file contain?
The JSON file contains a list of dictionaries, where each dictionary represents a predicted key-value relationship. Every entry includes the key text and bounding box coordinates, paired with the corresponding value text and its bounding box, enabling precise spatial and semantic extraction from the source document.
How does the visualization image represent predicted relationships?
The PNG visualization uses Pillow to draw lines connecting the center points of predicted key-value pairs. The key endpoint renders in green while the value endpoint renders in red, providing immediate visual feedback on the model's link predictions overlaid on the original document image.
Can I customize the output directory for inference results?
The current implementation in doc2graph/paths.py hardcodes the INFERENCE path as ROOT / "inference". To change the output location, you must modify the INFERENCE constant in doc2graph/paths.py or manually move files after generation, as the CLI and programmatic interfaces do not expose a configurable output path parameter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →