How to Debug Issues with Graph Construction from PDF Documents in doc2graph

When doc2graph receives a PDF input, it fails because the __fromPDF method in GraphBuilder is an unimplemented placeholder that returns None, causing downstream AttributeError exceptions or silent crashes during graph construction.

To debug PDF processing failures in andreagemelli/doc2graph, you must trace the execution from CLI arguments through the configuration system to the dispatcher in GraphBuilder.get_graph(), identify the missing implementation branch, and either implement the conversion logic or verify your input paths. The repository currently supports images and FUNSD datasets natively, but requires manual extension to handle PDF documents.

Understanding the PDF Processing Architecture

The path from a PDF file to a DGL graph follows a specific dispatch chain that terminates in an empty method.

  1. CLI Parsing: Arguments --src-data CUSTOM and --data-type pdf are parsed in doc2graph/main.py (lines 49-62) and stored in the runtime configuration.
  2. Configuration Loading: The get_config function in doc2graph/utils.py (lines 48-66) loads configs/preprocessing.yaml, making data_type available to the pipeline.
  3. GraphBuilder Dispatch: The GraphBuilder class in doc2graph/data/graph_builder.py receives the configuration. Its get_graph() method (lines 40-51) checks self.data_type and routes PDF requests to __fromPDF().
  4. The Placeholder: The __fromPDF method (lines 64-67) contains only a TODO comment and implicitly returns None, preventing any graph construction.

Because this method returns None, any downstream code expecting a tuple of (graphs, node_labels, edge_labels, features) will fail when attempting to access attributes or iterate over the result.

Common Debugging Symptoms and Root Causes

Identifying where the pipeline breaks depends on recognizing these specific failure modes:

  • AttributeError: 'NoneType' object has no attribute 'edate': The training or inference script received None from get_graph() and attempted to access graph properties. This confirms the __fromPDF placeholder was executed.
  • Silent termination with no output: The script parses arguments but exits after preprocessing without building graphs. Check if __fromPDF is called but returns nothing.
  • FileNotFoundError for PDF paths: The src_path passed to get_graph() does not exist, or the system lacks poppler-utils required by pdf2image.
  • Empty node and edge lists: If you implement a custom loader but see zero nodes, the PDF pages are not being rasterized correctly, or EasyOCR is returning empty detections on the converted images.

Step-by-Step Debugging Checklist

Follow this sequence to isolate the failure point:

  1. Verify CLI Flags: Run your command with verbose output and confirm args.data_type equals "pdf".

    python -m doc2graph.main --src-data CUSTOM --data-type pdf --edge-type knn
  2. Inspect the Configuration: Ensure the preprocessing config reflects the PDF data type.

    from doc2graph.utils import get_config
    cfg = get_config("preprocessing")
    print(cfg.GRAPHS.data_type)  # Must output: pdf
    
  3. Confirm Dispatch Execution: Add temporary logging inside GraphBuilder.get_graph() (around line 45) to verify the PDF branch is taken.

    elif self.data_type == "pdf":
        print(f"[DEBUG] Routing to __fromPDF with path: {src_path}")
        return self.__fromPDF(src_path)
  4. Check Method Return Values: If execution reaches __fromPDF, modify the method to print its inputs and confirm it is receiving a valid file path before implementing the conversion logic.

  5. Validate File Accessibility: Ensure the PDF exists at the specified path and that the process has read permissions.

Implementing the Missing PDF Pipeline

To resolve the None return issue, implement __fromPDF to convert PDF pages to temporary images, then reuse the existing __fromIMG method which contains the full OCR and graph construction logic.

Insert this implementation into doc2graph/data/graph_builder.py, replacing the placeholder at lines 64-67:

from pdf2image import convert_from_path
import tempfile
import shutil
import os
from typing import Tuple, List
import dgl  # Assuming DGL is imported elsewhere in the file

def __fromPDF(self, src_path: str) -> Tuple[List, List, List, List]:
    """
    Convert a PDF to DGL graphs by rasterizing pages and processing 
    them through the image pipeline.
    """
    # 1. Convert PDF pages to PNG images (requires poppler)

    pages = convert_from_path(src_path, fmt="png", dpi=300)
    
    # 2. Create temporary directory for intermediate files

    tmp_dir = tempfile.mkdtemp()
    img_paths = []
    
    try:
        for i, page in enumerate(pages):
            img_path = os.path.join(tmp_dir, f"page_{i}.png")
            page.save(img_path, "PNG")
            img_paths.append(img_path)
        
        # 3. Reuse existing image graph construction logic

        graphs, node_labels, edge_labels, features = self.__fromIMG(img_paths)
        
    finally:
        # 4. Cleanup temporary files regardless of success/failure

        shutil.rmtree(tmp_dir)
    
    return graphs, node_labels, edge_labels, features

Key Implementation Details:

  • convert_from_path: Renders each PDF page at 300 DPI to ensure OCR accuracy. This requires the poppler-utils system package on Linux or Poppler on Windows/macOS.
  • tempfile.mkdtemp: Isolates intermediate PNGs to prevent filesystem clutter.
  • self.__fromIMG: Leverages the existing EasyOCR integration and DGL graph builder, ensuring consistency with image-based workflows.
  • shutil.rmtree: Prevents disk space exhaustion when processing large PDF volumes.

Install the required dependency and system library:

pip install pdf2image

# Ubuntu/Debian:

sudo apt-get install poppler-utils

# macOS:

brew install poppler

Key Source Files for Debugging

Reference these specific locations when tracing execution or adding instrumentation:

File Purpose Critical Lines
doc2graph/main.py CLI entry point; defines --data-type argument 49-62
doc2graph/utils.py Configuration loader (get_config) 48-66
doc2graph/data/graph_builder.py GraphBuilder class and dispatch logic 19-28 (class), 40-51 (get_graph), 64-67 (__fromPDF placeholder)
configs/preprocessing.yaml Runtime settings for graph construction GRAPHS.data_type field

Summary

  • The Root Cause: doc2graph fails on PDFs because GraphBuilder.__fromPDF() is an unimplemented placeholder returning None.
  • The Symptom: Downstream code throws AttributeError when accessing graph properties on the None return value.
  • The Fix: Implement __fromPDF to convert PDF pages to temporary images using pdf2image, then pass those images to the existing __fromIMG method.
  • The Validation: After implementation, verify that get_graph() returns a non-empty list where each graph has num_nodes > 0 and valid feature dictionaries.
  • Dependencies: You must install pdf2image and the system poppler utilities to enable PDF rasterization.

Frequently Asked Questions

Why does doc2graph return NoneType errors when I use --data-type pdf?

The GraphBuilder class dispatches PDF inputs to the __fromPDF method, which currently contains only a TODO comment and returns None implicitly. When the training script attempts to access graph attributes like edate or num_nodes on this None value, it raises an AttributeError. You must implement the __fromPDF method to return valid graph tuples.

Can I process PDFs without modifying the doc2graph source code?

No. As of the current implementation, there is no hook or plugin system to inject custom PDF parsers. You must edit doc2graph/data/graph_builder.py to replace the placeholder __fromPDF method with working code that converts PDFs to a format the pipeline can process, typically by rasterizing pages to images.

What OCR engine processes text in PDF documents?

The repository uses EasyOCR for text detection and recognition. When you implement the PDF pipeline by converting pages to temporary PNG images, the existing __fromIMG method automatically applies EasyOCR to extract bounding boxes and text content, then constructs the graph using those detections.

How do I verify that my PDF graphs are constructed correctly?

After implementing __fromPDF, run a single-page test and inspect the returned objects:

builder = GraphBuilder()
graphs, _, _, feats = builder.get_graph("test.pdf", "CUSTOM")
assert graphs is not None, "Graph construction returned None"
assert len(graphs) > 0, "Empty graph list"
assert graphs[0].num_nodes() > 0, "Graph has no nodes"
print(f"Successfully built graph with {graphs[0].num_nodes()} nodes")

If all assertions pass and feats["boxs"] contains coordinate arrays, the PDF pipeline is functioning correctly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →