How to Switch Between FUNSD and PAU Datasets for Training in Doc2Graph
Use the --src-data command-line argument in doc2graph/main.py to select either "FUNSD" or "PAU", which routes training to the corresponding dataset-specific pipeline.
Doc2Graph supports two built-in document-understanding datasets—FUNSD and PAU—and allows you to switch between them at runtime without modifying the source code. By passing a single argument to the training script, you control which data loader and training routine the framework invokes. This guide explains the mechanism behind dataset selection and provides practical commands for both CLI and programmatic usage.
How Dataset Selection Works in Doc2Graph
The dataset switch is handled by the argument parser in doc2graph/main.py. The --src-data parameter accepts a string value that determines which training pipeline executes.
Argument Definition and Defaults
In doc2graph/main.py, the parser defines --src-data with a default value of "FUNSD":
# Lines 49-55 in doc2graph/main.py
parser.add_argument(
"--src-data",
type=str,
default="FUNSD",
help="Dataset to use: FUNSD or PAU"
)
Runtime Routing Logic
After parsing, the script branches on args.src_data (lines 29-45 in doc2graph/main.py):
- If the value is
"FUNSD", the code callstrain_funsd(args)to launch the FUNSD-specific training routine. - If the value is
"PAU", the code callstrain_pau(args)to launch the PAU-specific training routine. - Any other value raises an exception, enforcing strict validation of supported datasets.
This branching ensures that only the supported datasets are used and that the correct data paths and preprocessing logic are loaded automatically.
Dataset Paths and Initialization
Before training, Doc2Graph needs to know where the datasets reside. These locations are defined centrally in doc2graph/paths.py.
Centralized Path Configuration
The file doc2graph/paths.py (lines 23-30) declares four key constants:
FUNSD_TRAINandFUNSD_TESTfor the FUNSD dataset splitsPAU_TRAINandPAU_TESTfor the PAU dataset splits
These paths point to the expected folder layout under the top-level DATA directory.
Automatic Dataset Download
You do not need to manually download or organize the datasets. The --init flag in main.py handles dataset preparation:
uv run python -m doc2graph.main --init
This command downloads both FUNSD and PAU datasets and creates the necessary folder structure under DATA, ensuring FUNSD_TRAIN, FUNSD_TEST, PAU_TRAIN, and PAU_TEST resolve correctly.
Training Commands for Each Dataset
Once initialized, switch between datasets by changing the --src-data value. The rest of your configuration—model architecture, GPU selection, and feature flags—remains identical.
Training on FUNSD (Default)
Since FUNSD is the default, you can omit --src-data or include it for clarity:
uv run python -m doc2graph.main \
--src-data FUNSD \
--model e2e \
--gpu 0 \
-addG -addT -addE -addV
Training on PAU
To switch to the PAU dataset, simply change the argument value:
uv run python -m doc2graph.main \
--src-data PAU \
--model e2e \
--gpu 0 \
-addG -addT -addE -addV
According to the doc2graph source code, the train_pau(args) function in doc2graph/training/pau.py (lines 16-18) uses the PAU_TRAIN and PAU_TEST paths defined in paths.py, while train_funsd(args) in doc2graph/training/funsd.py (lines 13-18) uses the FUNSD equivalents.
Programmatic Dataset Switching
You can also invoke the training pipeline from Python scripts or notebooks by simulating command-line arguments:
from doc2graph.main import main
import sys
# Configure arguments for PAU training
sys.argv = [
"doc2graph.main",
"--src-data", "PAU",
"--model", "e2e",
"--gpu", "0",
"-addG", "-addT", "-addE", "-addV"
]
# Execute the training pipeline
main()
This approach is useful for hyperparameter sweeps or integration with experiment tracking tools.
Summary
- Use
--src-datafollowed by"FUNSD"or"PAU"to select your dataset indoc2graph/main.py. - FUNSD is the default when the argument is omitted, routing to
train_funsd(args)indoc2graph/training/funsd.py. - PAU training invokes
train_pau(args)fromdoc2graph/training/pau.pyusing paths defined indoc2graph/paths.py. - Initialize once with
--initto download both datasets; switching requires only changing the--src-datavalue. - Programmatic control is available by setting
sys.argvbefore callingmain().
Frequently Asked Questions
What happens if I pass an unsupported dataset name to --src-data?
The script validates the args.src_data value in doc2graph/main.py and raises an exception if it is neither "FUNSD" nor "PAU". This prevents execution with invalid data paths and ensures only supported datasets are processed.
Do I need to reinitialize the data when switching between FUNSD and PAU?
No. The --init flag downloads and prepares both datasets simultaneously. Once initialized, you can switch between them arbitrarily by changing the --src-data argument without re-downloading or reprocessing.
Can I use custom dataset paths instead of the defaults?
The training routines in doc2graph/training/funsd.py and doc2graph/training/pau.py import paths from doc2graph/paths.py. To use custom locations, modify the constants (FUNSD_TRAIN, PAU_TRAIN, etc.) in paths.py, or override the path variables before importing the training modules.
Is the model architecture different between FUNSD and PAU training?
No. The --model argument and feature flags (-addG, -addT, -addE, -addV) apply identically to both datasets. The only difference is the data loader and training routine invoked based on the --src-data selection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →