# How to Debug Issues with Graph Construction from PDF Documents in doc2graph

> Debug doc2graph PDF issues. Learn how to fix unimplemented methods causing errors during graph construction from PDF files and prevent crashes.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**When `doc2graph` receives a PDF input, it fails because the `__fromPDF` method in `GraphBuilder` is an unimplemented placeholder that returns `None`, causing downstream `AttributeError` exceptions or silent crashes during graph construction.**

To debug PDF processing failures in `andreagemelli/doc2graph`, you must trace the execution from CLI arguments through the configuration system to the dispatcher in `GraphBuilder.get_graph()`, identify the missing implementation branch, and either implement the conversion logic or verify your input paths. The repository currently supports images and FUNSD datasets natively, but requires manual extension to handle PDF documents.

## Understanding the PDF Processing Architecture

The path from a PDF file to a DGL graph follows a specific dispatch chain that terminates in an empty method.

1.  **CLI Parsing**: Arguments `--src-data CUSTOM` and `--data-type pdf` are parsed in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) (lines 49-62) and stored in the runtime configuration.
2.  **Configuration Loading**: The `get_config` function in [`doc2graph/utils.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/utils.py) (lines 48-66) loads [`configs/preprocessing.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/configs/preprocessing.yaml), making `data_type` available to the pipeline.
3.  **GraphBuilder Dispatch**: The `GraphBuilder` class in [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) receives the configuration. Its `get_graph()` method (lines 40-51) checks `self.data_type` and routes PDF requests to `__fromPDF()`.
4.  **The Placeholder**: The `__fromPDF` method (lines 64-67) contains only a `TODO` comment and implicitly returns `None`, preventing any graph construction.

Because this method returns `None`, any downstream code expecting a tuple of `(graphs, node_labels, edge_labels, features)` will fail when attempting to access attributes or iterate over the result.

## Common Debugging Symptoms and Root Causes

Identifying where the pipeline breaks depends on recognizing these specific failure modes:

*   **`AttributeError: 'NoneType' object has no attribute 'edate'`**: The training or inference script received `None` from `get_graph()` and attempted to access graph properties. This confirms the `__fromPDF` placeholder was executed.
*   **Silent termination with no output**: The script parses arguments but exits after preprocessing without building graphs. Check if `__fromPDF` is called but returns nothing.
*   **`FileNotFoundError` for PDF paths**: The `src_path` passed to `get_graph()` does not exist, or the system lacks `poppler-utils` required by `pdf2image`.
*   **Empty node and edge lists**: If you implement a custom loader but see zero nodes, the PDF pages are not being rasterized correctly, or EasyOCR is returning empty detections on the converted images.

## Step-by-Step Debugging Checklist

Follow this sequence to isolate the failure point:

1.  **Verify CLI Flags**: Run your command with verbose output and confirm `args.data_type` equals `"pdf"`.
    ```bash
    python -m doc2graph.main --src-data CUSTOM --data-type pdf --edge-type knn
    ```

2.  **Inspect the Configuration**: Ensure the preprocessing config reflects the PDF data type.
    ```python
    from doc2graph.utils import get_config
    cfg = get_config("preprocessing")
    print(cfg.GRAPHS.data_type)  # Must output: pdf

    ```

3.  **Confirm Dispatch Execution**: Add temporary logging inside `GraphBuilder.get_graph()` (around line 45) to verify the PDF branch is taken.
    ```python
    elif self.data_type == "pdf":
        print(f"[DEBUG] Routing to __fromPDF with path: {src_path}")
        return self.__fromPDF(src_path)
    ```

4.  **Check Method Return Values**: If execution reaches `__fromPDF`, modify the method to print its inputs and confirm it is receiving a valid file path before implementing the conversion logic.

5.  **Validate File Accessibility**: Ensure the PDF exists at the specified path and that the process has read permissions.

## Implementing the Missing PDF Pipeline

To resolve the `None` return issue, implement `__fromPDF` to convert PDF pages to temporary images, then reuse the existing `__fromIMG` method which contains the full OCR and graph construction logic.

Insert this implementation into [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py), replacing the placeholder at lines 64-67:

```python
from pdf2image import convert_from_path
import tempfile
import shutil
import os
from typing import Tuple, List
import dgl  # Assuming DGL is imported elsewhere in the file

def __fromPDF(self, src_path: str) -> Tuple[List, List, List, List]:
    """
    Convert a PDF to DGL graphs by rasterizing pages and processing 
    them through the image pipeline.
    """
    # 1. Convert PDF pages to PNG images (requires poppler)

    pages = convert_from_path(src_path, fmt="png", dpi=300)
    
    # 2. Create temporary directory for intermediate files

    tmp_dir = tempfile.mkdtemp()
    img_paths = []
    
    try:
        for i, page in enumerate(pages):
            img_path = os.path.join(tmp_dir, f"page_{i}.png")
            page.save(img_path, "PNG")
            img_paths.append(img_path)
        
        # 3. Reuse existing image graph construction logic

        graphs, node_labels, edge_labels, features = self.__fromIMG(img_paths)
        
    finally:
        # 4. Cleanup temporary files regardless of success/failure

        shutil.rmtree(tmp_dir)
    
    return graphs, node_labels, edge_labels, features

```

**Key Implementation Details**:

*   **`convert_from_path`**: Renders each PDF page at 300 DPI to ensure OCR accuracy. This requires the `poppler-utils` system package on Linux or Poppler on Windows/macOS.
*   **`tempfile.mkdtemp`**: Isolates intermediate PNGs to prevent filesystem clutter.
*   **`self.__fromIMG`**: Leverages the existing EasyOCR integration and DGL graph builder, ensuring consistency with image-based workflows.
*   **`shutil.rmtree`**: Prevents disk space exhaustion when processing large PDF volumes.

Install the required dependency and system library:

```bash
pip install pdf2image

# Ubuntu/Debian:

sudo apt-get install poppler-utils

# macOS:

brew install poppler

```

## Key Source Files for Debugging

Reference these specific locations when tracing execution or adding instrumentation:

| File | Purpose | Critical Lines |
|------|---------|----------------|
| [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) | CLI entry point; defines `--data-type` argument | 49-62 |
| [`doc2graph/utils.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/utils.py) | Configuration loader (`get_config`) | 48-66 |
| [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) | `GraphBuilder` class and dispatch logic | 19-28 (class), 40-51 (`get_graph`), 64-67 (`__fromPDF` placeholder) |
| [`configs/preprocessing.yaml`](https://github.com/andreagemelli/doc2graph/blob/main/configs/preprocessing.yaml) | Runtime settings for graph construction | `GRAPHS.data_type` field |

## Summary

*   **The Root Cause**: `doc2graph` fails on PDFs because `GraphBuilder.__fromPDF()` is an unimplemented placeholder returning `None`.
*   **The Symptom**: Downstream code throws `AttributeError` when accessing graph properties on the `None` return value.
*   **The Fix**: Implement `__fromPDF` to convert PDF pages to temporary images using `pdf2image`, then pass those images to the existing `__fromIMG` method.
*   **The Validation**: After implementation, verify that `get_graph()` returns a non-empty list where each graph has `num_nodes > 0` and valid feature dictionaries.
*   **Dependencies**: You must install `pdf2image` and the system `poppler` utilities to enable PDF rasterization.

## Frequently Asked Questions

### Why does doc2graph return `NoneType` errors when I use `--data-type pdf`?

The `GraphBuilder` class dispatches PDF inputs to the `__fromPDF` method, which currently contains only a `TODO` comment and returns `None` implicitly. When the training script attempts to access graph attributes like `edate` or `num_nodes` on this `None` value, it raises an `AttributeError`. You must implement the `__fromPDF` method to return valid graph tuples.

### Can I process PDFs without modifying the doc2graph source code?

No. As of the current implementation, there is no hook or plugin system to inject custom PDF parsers. You must edit [`doc2graph/data/graph_builder.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/data/graph_builder.py) to replace the placeholder `__fromPDF` method with working code that converts PDFs to a format the pipeline can process, typically by rasterizing pages to images.

### What OCR engine processes text in PDF documents?

The repository uses **EasyOCR** for text detection and recognition. When you implement the PDF pipeline by converting pages to temporary PNG images, the existing `__fromIMG` method automatically applies EasyOCR to extract bounding boxes and text content, then constructs the graph using those detections.

### How do I verify that my PDF graphs are constructed correctly?

After implementing `__fromPDF`, run a single-page test and inspect the returned objects:

```python
builder = GraphBuilder()
graphs, _, _, feats = builder.get_graph("test.pdf", "CUSTOM")
assert graphs is not None, "Graph construction returned None"
assert len(graphs) > 0, "Empty graph list"
assert graphs[0].num_nodes() > 0, "Graph has no nodes"
print(f"Successfully built graph with {graphs[0].num_nodes()} nodes")

```

If all assertions pass and `feats["boxs"]` contains coordinate arrays, the PDF pipeline is functioning correctly.