How to Debug Issues with Graph Construction from PDF Documents in doc2graph
When doc2graph receives a PDF input, it fails because the __fromPDF method in GraphBuilder is an unimplemented placeholder that returns None, causing downstream AttributeError exceptions or silent crashes during graph construction.
To debug PDF processing failures in andreagemelli/doc2graph, you must trace the execution from CLI arguments through the configuration system to the dispatcher in GraphBuilder.get_graph(), identify the missing implementation branch, and either implement the conversion logic or verify your input paths. The repository currently supports images and FUNSD datasets natively, but requires manual extension to handle PDF documents.
Understanding the PDF Processing Architecture
The path from a PDF file to a DGL graph follows a specific dispatch chain that terminates in an empty method.
- CLI Parsing: Arguments
--src-data CUSTOMand--data-type pdfare parsed indoc2graph/main.py(lines 49-62) and stored in the runtime configuration. - Configuration Loading: The
get_configfunction indoc2graph/utils.py(lines 48-66) loadsconfigs/preprocessing.yaml, makingdata_typeavailable to the pipeline. - GraphBuilder Dispatch: The
GraphBuilderclass indoc2graph/data/graph_builder.pyreceives the configuration. Itsget_graph()method (lines 40-51) checksself.data_typeand routes PDF requests to__fromPDF(). - The Placeholder: The
__fromPDFmethod (lines 64-67) contains only aTODOcomment and implicitly returnsNone, preventing any graph construction.
Because this method returns None, any downstream code expecting a tuple of (graphs, node_labels, edge_labels, features) will fail when attempting to access attributes or iterate over the result.
Common Debugging Symptoms and Root Causes
Identifying where the pipeline breaks depends on recognizing these specific failure modes:
AttributeError: 'NoneType' object has no attribute 'edate': The training or inference script receivedNonefromget_graph()and attempted to access graph properties. This confirms the__fromPDFplaceholder was executed.- Silent termination with no output: The script parses arguments but exits after preprocessing without building graphs. Check if
__fromPDFis called but returns nothing. FileNotFoundErrorfor PDF paths: Thesrc_pathpassed toget_graph()does not exist, or the system lackspoppler-utilsrequired bypdf2image.- Empty node and edge lists: If you implement a custom loader but see zero nodes, the PDF pages are not being rasterized correctly, or EasyOCR is returning empty detections on the converted images.
Step-by-Step Debugging Checklist
Follow this sequence to isolate the failure point:
-
Verify CLI Flags: Run your command with verbose output and confirm
args.data_typeequals"pdf".python -m doc2graph.main --src-data CUSTOM --data-type pdf --edge-type knn -
Inspect the Configuration: Ensure the preprocessing config reflects the PDF data type.
from doc2graph.utils import get_config cfg = get_config("preprocessing") print(cfg.GRAPHS.data_type) # Must output: pdf -
Confirm Dispatch Execution: Add temporary logging inside
GraphBuilder.get_graph()(around line 45) to verify the PDF branch is taken.elif self.data_type == "pdf": print(f"[DEBUG] Routing to __fromPDF with path: {src_path}") return self.__fromPDF(src_path) -
Check Method Return Values: If execution reaches
__fromPDF, modify the method to print its inputs and confirm it is receiving a valid file path before implementing the conversion logic. -
Validate File Accessibility: Ensure the PDF exists at the specified path and that the process has read permissions.
Implementing the Missing PDF Pipeline
To resolve the None return issue, implement __fromPDF to convert PDF pages to temporary images, then reuse the existing __fromIMG method which contains the full OCR and graph construction logic.
Insert this implementation into doc2graph/data/graph_builder.py, replacing the placeholder at lines 64-67:
from pdf2image import convert_from_path
import tempfile
import shutil
import os
from typing import Tuple, List
import dgl # Assuming DGL is imported elsewhere in the file
def __fromPDF(self, src_path: str) -> Tuple[List, List, List, List]:
"""
Convert a PDF to DGL graphs by rasterizing pages and processing
them through the image pipeline.
"""
# 1. Convert PDF pages to PNG images (requires poppler)
pages = convert_from_path(src_path, fmt="png", dpi=300)
# 2. Create temporary directory for intermediate files
tmp_dir = tempfile.mkdtemp()
img_paths = []
try:
for i, page in enumerate(pages):
img_path = os.path.join(tmp_dir, f"page_{i}.png")
page.save(img_path, "PNG")
img_paths.append(img_path)
# 3. Reuse existing image graph construction logic
graphs, node_labels, edge_labels, features = self.__fromIMG(img_paths)
finally:
# 4. Cleanup temporary files regardless of success/failure
shutil.rmtree(tmp_dir)
return graphs, node_labels, edge_labels, features
Key Implementation Details:
convert_from_path: Renders each PDF page at 300 DPI to ensure OCR accuracy. This requires thepoppler-utilssystem package on Linux or Poppler on Windows/macOS.tempfile.mkdtemp: Isolates intermediate PNGs to prevent filesystem clutter.self.__fromIMG: Leverages the existing EasyOCR integration and DGL graph builder, ensuring consistency with image-based workflows.shutil.rmtree: Prevents disk space exhaustion when processing large PDF volumes.
Install the required dependency and system library:
pip install pdf2image
# Ubuntu/Debian:
sudo apt-get install poppler-utils
# macOS:
brew install poppler
Key Source Files for Debugging
Reference these specific locations when tracing execution or adding instrumentation:
| File | Purpose | Critical Lines |
|---|---|---|
doc2graph/main.py |
CLI entry point; defines --data-type argument |
49-62 |
doc2graph/utils.py |
Configuration loader (get_config) |
48-66 |
doc2graph/data/graph_builder.py |
GraphBuilder class and dispatch logic |
19-28 (class), 40-51 (get_graph), 64-67 (__fromPDF placeholder) |
configs/preprocessing.yaml |
Runtime settings for graph construction | GRAPHS.data_type field |
Summary
- The Root Cause:
doc2graphfails on PDFs becauseGraphBuilder.__fromPDF()is an unimplemented placeholder returningNone. - The Symptom: Downstream code throws
AttributeErrorwhen accessing graph properties on theNonereturn value. - The Fix: Implement
__fromPDFto convert PDF pages to temporary images usingpdf2image, then pass those images to the existing__fromIMGmethod. - The Validation: After implementation, verify that
get_graph()returns a non-empty list where each graph hasnum_nodes > 0and valid feature dictionaries. - Dependencies: You must install
pdf2imageand the systempopplerutilities to enable PDF rasterization.
Frequently Asked Questions
Why does doc2graph return NoneType errors when I use --data-type pdf?
The GraphBuilder class dispatches PDF inputs to the __fromPDF method, which currently contains only a TODO comment and returns None implicitly. When the training script attempts to access graph attributes like edate or num_nodes on this None value, it raises an AttributeError. You must implement the __fromPDF method to return valid graph tuples.
Can I process PDFs without modifying the doc2graph source code?
No. As of the current implementation, there is no hook or plugin system to inject custom PDF parsers. You must edit doc2graph/data/graph_builder.py to replace the placeholder __fromPDF method with working code that converts PDFs to a format the pipeline can process, typically by rasterizing pages to images.
What OCR engine processes text in PDF documents?
The repository uses EasyOCR for text detection and recognition. When you implement the PDF pipeline by converting pages to temporary PNG images, the existing __fromIMG method automatically applies EasyOCR to extract bounding boxes and text content, then constructs the graph using those detections.
How do I verify that my PDF graphs are constructed correctly?
After implementing __fromPDF, run a single-page test and inspect the returned objects:
builder = GraphBuilder()
graphs, _, _, feats = builder.get_graph("test.pdf", "CUSTOM")
assert graphs is not None, "Graph construction returned None"
assert len(graphs) > 0, "Empty graph list"
assert graphs[0].num_nodes() > 0, "Graph has no nodes"
print(f"Successfully built graph with {graphs[0].num_nodes()} nodes")
If all assertions pass and feats["boxs"] contains coordinate arrays, the PDF pipeline is functioning correctly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →