How to Switch Between FUNSD and PAU Datasets for Training in Doc2Graph

Use the --src-data command-line argument in doc2graph/main.py to select either "FUNSD" or "PAU", which routes training to the corresponding dataset-specific pipeline.

Doc2Graph supports two built-in document-understanding datasets—FUNSD and PAU—and allows you to switch between them at runtime without modifying the source code. By passing a single argument to the training script, you control which data loader and training routine the framework invokes. This guide explains the mechanism behind dataset selection and provides practical commands for both CLI and programmatic usage.

How Dataset Selection Works in Doc2Graph

The dataset switch is handled by the argument parser in doc2graph/main.py. The --src-data parameter accepts a string value that determines which training pipeline executes.

Argument Definition and Defaults

In doc2graph/main.py, the parser defines --src-data with a default value of "FUNSD":


# Lines 49-55 in doc2graph/main.py

parser.add_argument(
    "--src-data",
    type=str,
    default="FUNSD",
    help="Dataset to use: FUNSD or PAU"
)

Runtime Routing Logic

After parsing, the script branches on args.src_data (lines 29-45 in doc2graph/main.py):

  • If the value is "FUNSD", the code calls train_funsd(args) to launch the FUNSD-specific training routine.
  • If the value is "PAU", the code calls train_pau(args) to launch the PAU-specific training routine.
  • Any other value raises an exception, enforcing strict validation of supported datasets.

This branching ensures that only the supported datasets are used and that the correct data paths and preprocessing logic are loaded automatically.

Dataset Paths and Initialization

Before training, Doc2Graph needs to know where the datasets reside. These locations are defined centrally in doc2graph/paths.py.

Centralized Path Configuration

The file doc2graph/paths.py (lines 23-30) declares four key constants:

  • FUNSD_TRAIN and FUNSD_TEST for the FUNSD dataset splits
  • PAU_TRAIN and PAU_TEST for the PAU dataset splits

These paths point to the expected folder layout under the top-level DATA directory.

Automatic Dataset Download

You do not need to manually download or organize the datasets. The --init flag in main.py handles dataset preparation:

uv run python -m doc2graph.main --init

This command downloads both FUNSD and PAU datasets and creates the necessary folder structure under DATA, ensuring FUNSD_TRAIN, FUNSD_TEST, PAU_TRAIN, and PAU_TEST resolve correctly.

Training Commands for Each Dataset

Once initialized, switch between datasets by changing the --src-data value. The rest of your configuration—model architecture, GPU selection, and feature flags—remains identical.

Training on FUNSD (Default)

Since FUNSD is the default, you can omit --src-data or include it for clarity:

uv run python -m doc2graph.main \
    --src-data FUNSD \
    --model e2e \
    --gpu 0 \
    -addG -addT -addE -addV

Training on PAU

To switch to the PAU dataset, simply change the argument value:

uv run python -m doc2graph.main \
    --src-data PAU \
    --model e2e \
    --gpu 0 \
    -addG -addT -addE -addV

According to the doc2graph source code, the train_pau(args) function in doc2graph/training/pau.py (lines 16-18) uses the PAU_TRAIN and PAU_TEST paths defined in paths.py, while train_funsd(args) in doc2graph/training/funsd.py (lines 13-18) uses the FUNSD equivalents.

Programmatic Dataset Switching

You can also invoke the training pipeline from Python scripts or notebooks by simulating command-line arguments:

from doc2graph.main import main
import sys

# Configure arguments for PAU training

sys.argv = [
    "doc2graph.main",
    "--src-data", "PAU",
    "--model", "e2e",
    "--gpu", "0",
    "-addG", "-addT", "-addE", "-addV"
]

# Execute the training pipeline

main()

This approach is useful for hyperparameter sweeps or integration with experiment tracking tools.

Summary

  • Use --src-data followed by "FUNSD" or "PAU" to select your dataset in doc2graph/main.py.
  • FUNSD is the default when the argument is omitted, routing to train_funsd(args) in doc2graph/training/funsd.py.
  • PAU training invokes train_pau(args) from doc2graph/training/pau.py using paths defined in doc2graph/paths.py.
  • Initialize once with --init to download both datasets; switching requires only changing the --src-data value.
  • Programmatic control is available by setting sys.argv before calling main().

Frequently Asked Questions

What happens if I pass an unsupported dataset name to --src-data?

The script validates the args.src_data value in doc2graph/main.py and raises an exception if it is neither "FUNSD" nor "PAU". This prevents execution with invalid data paths and ensures only supported datasets are processed.

Do I need to reinitialize the data when switching between FUNSD and PAU?

No. The --init flag downloads and prepares both datasets simultaneously. Once initialized, you can switch between them arbitrarily by changing the --src-data argument without re-downloading or reprocessing.

Can I use custom dataset paths instead of the defaults?

The training routines in doc2graph/training/funsd.py and doc2graph/training/pau.py import paths from doc2graph/paths.py. To use custom locations, modify the constants (FUNSD_TRAIN, PAU_TRAIN, etc.) in paths.py, or override the path variables before importing the training modules.

Is the model architecture different between FUNSD and PAU training?

No. The --model argument and feature flags (-addG, -addT, -addE, -addV) apply identically to both datasets. The only difference is the data loader and training routine invoked based on the --src-data selection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →