# How to Switch Between FUNSD and PAU Datasets for Training in Doc2Graph

> Easily switch between FUNSD and PAU datasets for Doc2Graph training using the `src-data` argument in main.py. Select your preferred dataset and optimize your graph generation.

- Repository: [Andrea Gemelli/doc2graph](https://github.com/andreagemelli/doc2graph)
- Tags: how-to-guide
- Published: 2026-02-24

---

**Use the `--src-data` command-line argument in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) to select either "FUNSD" or "PAU", which routes training to the corresponding dataset-specific pipeline.**

Doc2Graph supports two built-in document-understanding datasets—**FUNSD** and **PAU**—and allows you to switch between them at runtime without modifying the source code. By passing a single argument to the training script, you control which data loader and training routine the framework invokes. This guide explains the mechanism behind dataset selection and provides practical commands for both CLI and programmatic usage.

## How Dataset Selection Works in Doc2Graph

The dataset switch is handled by the argument parser in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py). The `--src-data` parameter accepts a string value that determines which training pipeline executes.

### Argument Definition and Defaults

In [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py), the parser defines `--src-data` with a default value of `"FUNSD"`:

```python

# Lines 49-55 in doc2graph/main.py

parser.add_argument(
    "--src-data",
    type=str,
    default="FUNSD",
    help="Dataset to use: FUNSD or PAU"
)

```

### Runtime Routing Logic

After parsing, the script branches on `args.src_data` (lines 29-45 in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py)):

- If the value is `"FUNSD"`, the code calls `train_funsd(args)` to launch the FUNSD-specific training routine.
- If the value is `"PAU"`, the code calls `train_pau(args)` to launch the PAU-specific training routine.
- Any other value raises an exception, enforcing strict validation of supported datasets.

This branching ensures that only the supported datasets are used and that the correct data paths and preprocessing logic are loaded automatically.

## Dataset Paths and Initialization

Before training, Doc2Graph needs to know where the datasets reside. These locations are defined centrally in [`doc2graph/paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/paths.py).

### Centralized Path Configuration

The file [`doc2graph/paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/paths.py) (lines 23-30) declares four key constants:
- `FUNSD_TRAIN` and `FUNSD_TEST` for the FUNSD dataset splits
- `PAU_TRAIN` and `PAU_TEST` for the PAU dataset splits

These paths point to the expected folder layout under the top-level `DATA` directory.

### Automatic Dataset Download

You do not need to manually download or organize the datasets. The `--init` flag in [`main.py`](https://github.com/andreagemelli/doc2graph/blob/main/main.py) handles dataset preparation:

```bash
uv run python -m doc2graph.main --init

```

This command downloads both FUNSD and PAU datasets and creates the necessary folder structure under `DATA`, ensuring `FUNSD_TRAIN`, `FUNSD_TEST`, `PAU_TRAIN`, and `PAU_TEST` resolve correctly.

## Training Commands for Each Dataset

Once initialized, switch between datasets by changing the `--src-data` value. The rest of your configuration—model architecture, GPU selection, and feature flags—remains identical.

### Training on FUNSD (Default)

Since FUNSD is the default, you can omit `--src-data` or include it for clarity:

```bash
uv run python -m doc2graph.main \
    --src-data FUNSD \
    --model e2e \
    --gpu 0 \
    -addG -addT -addE -addV

```

### Training on PAU

To switch to the PAU dataset, simply change the argument value:

```bash
uv run python -m doc2graph.main \
    --src-data PAU \
    --model e2e \
    --gpu 0 \
    -addG -addT -addE -addV

```

According to the doc2graph source code, the `train_pau(args)` function in [`doc2graph/training/pau.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/pau.py) (lines 16-18) uses the `PAU_TRAIN` and `PAU_TEST` paths defined in [`paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/paths.py), while `train_funsd(args)` in [`doc2graph/training/funsd.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/funsd.py) (lines 13-18) uses the FUNSD equivalents.

## Programmatic Dataset Switching

You can also invoke the training pipeline from Python scripts or notebooks by simulating command-line arguments:

```python
from doc2graph.main import main
import sys

# Configure arguments for PAU training

sys.argv = [
    "doc2graph.main",
    "--src-data", "PAU",
    "--model", "e2e",
    "--gpu", "0",
    "-addG", "-addT", "-addE", "-addV"
]

# Execute the training pipeline

main()

```

This approach is useful for hyperparameter sweeps or integration with experiment tracking tools.

## Summary

- **Use `--src-data`** followed by `"FUNSD"` or `"PAU"` to select your dataset in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py).
- **FUNSD is the default** when the argument is omitted, routing to `train_funsd(args)` in [`doc2graph/training/funsd.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/funsd.py).
- **PAU training** invokes `train_pau(args)` from [`doc2graph/training/pau.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/pau.py) using paths defined in [`doc2graph/paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/paths.py).
- **Initialize once** with `--init` to download both datasets; switching requires only changing the `--src-data` value.
- **Programmatic control** is available by setting `sys.argv` before calling `main()`.

## Frequently Asked Questions

### What happens if I pass an unsupported dataset name to `--src-data`?

The script validates the `args.src_data` value in [`doc2graph/main.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/main.py) and raises an exception if it is neither `"FUNSD"` nor `"PAU"`. This prevents execution with invalid data paths and ensures only supported datasets are processed.

### Do I need to reinitialize the data when switching between FUNSD and PAU?

No. The `--init` flag downloads and prepares both datasets simultaneously. Once initialized, you can switch between them arbitrarily by changing the `--src-data` argument without re-downloading or reprocessing.

### Can I use custom dataset paths instead of the defaults?

The training routines in [`doc2graph/training/funsd.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/funsd.py) and [`doc2graph/training/pau.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/training/pau.py) import paths from [`doc2graph/paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/doc2graph/paths.py). To use custom locations, modify the constants (`FUNSD_TRAIN`, `PAU_TRAIN`, etc.) in [`paths.py`](https://github.com/andreagemelli/doc2graph/blob/main/paths.py), or override the path variables before importing the training modules.

### Is the model architecture different between FUNSD and PAU training?

No. The `--model` argument and feature flags (`-addG`, `-addT`, `-addE`, `-addV`) apply identically to both datasets. The only difference is the data loader and training routine invoked based on the `--src-data` selection.