How to Create a Custom Model Configuration in doc2graph by Modifying YAML Files

To create a custom model configuration in doc2graph, copy an existing YAML file from configs/models/, modify the architectural hyper-parameters, and load it via doc2graph.utils.get_config which converts the YAML into an attribute-accessible dictionary for the training pipeline.

The doc2graph library stores all model hyper-parameters in dedicated YAML files under the configs/models/ directory. Instead of hard-coding parameters in Python scripts, the framework uses these configuration files to define model architectures, dropout rates, hidden dimensions, and training flags. This YAML-based approach allows you to experiment with new architectures by simply editing text files without touching the source code.

Step-by-Step Guide to Creating Custom Model Configurations

Choose a Base Configuration Template

Start by selecting an existing configuration file from configs/models/ as your template. The repository provides several ready-made configurations including base.yaml, gcn.yaml, edge.yaml, and e2e.yaml. Each file contains structured hyper-parameters specific to different model architectures. For example, gcn.yaml defines parameters such as dropout, hidden_dim, and num_layers for graph convolutional networks. Referencing an existing template ensures you include all required keys that the model constructors expect.

Create and Edit Your YAML File

Create a new YAML file in configs/models/ by copying your chosen template. Use the following command to duplicate an existing configuration:

cp configs/models/gcn.yaml configs/models/my_custom.yaml

Open the new file and modify any key-value pairs to define your custom architecture. Common configuration fields include:

  • name – A string identifier used for logging and checkpoint naming.
  • dropout – Float value for the dropout probability (e.g., 0.1).
  • hidden_dim – Integer specifying the size of hidden layers (e.g., 128).
  • num_layers – Integer defining the network depth (e.g., 4).
  • attn – Boolean flag to enable attention mechanisms.
  • out_chunks – Integer for output segmentation (e.g., 150).
  • doProject – Boolean specific to edge-oriented models (see edge.yaml).

Example configs/models/my_custom.yaml:

name: MY_CUSTOM
dropout: 0.1
hidden_dim: 128
num_layers: 4
attn: true
out_chunks: 150

Load the Configuration in Python

The framework loads YAML configurations through the get_config function in doc2graph/utils.py. This utility reads the specified YAML file and returns an AttrDict object that supports attribute-style access (e.g., cfg.hidden_dim instead of cfg['hidden_dim']).

According to lines 57-65 of doc2graph/utils.py, the implementation recursively converts nested dictionaries into AttrDict instances:

def get_config(name: str) -> AttrDict:
    with open(CONFIGS / f"{name}.yaml") as fileobj:
        config_dict = yaml.safe_load(fileobj)
        config = AttrDict()
        for key, value in config_dict.items():
            if isinstance(value, dict):
                config[key] = AttrDict(value)
            else:
                config[key] = value
    return config

To load your custom configuration in a script:

from doc2graph.utils import get_config

cfg = get_config("my_custom")   # loads configs/models/my_custom.yaml

print(cfg.hidden_dim)           # → 128

print(cfg.name)                 # → MY_CUSTOM

Run Training with the Custom Model

Pass the configuration stem (filename without .yaml extension) to the training script using the --model CLI argument. The entry point in doc2graph/main.py forwards this flag to get_config and initializes the model with your specified hyper-parameters:

python -m doc2graph.main --model my_custom

Configuration File Structure and Key Parameters

The configs/models/ directory serves as the central repository for all model definitions. Each YAML file follows a flat or shallowly-nested structure where top-level keys map directly to constructor arguments in the model classes. When defining custom model configurations, ensure you include architecture-specific keys:

  • For GCN models: hidden_dim, num_layers, dropout, and attn.
  • For edge-oriented models: Additional flags like doProject control projection layers before edge classification.

The name field is particularly important as it determines the subdirectory for saved checkpoints and log files during training runs.

After conducting hyper-parameter optimization, you may want to save the best configuration permanently. The training utilities in doc2graph/training/utils.py demonstrate this pattern using shutil.copyfile (lines 120-121), which copies the base template to a new file with finalized values. You can replicate this behavior to archive tuned configurations:

import shutil
from pathlib import Path

shutil.copyfile(
    Path("configs/models/base.yaml"), 
    Path("configs/models/best_tuned.yaml")
)

Summary

  • Custom model configurations in doc2graph reside in configs/models/ as YAML files.
  • Copy an existing template (such as gcn.yaml or edge.yaml) to create a new configuration without syntax errors.
  • Edit architectural parameters like hidden_dim, num_layers, and dropout to define model behavior.
  • Use doc2graph.utils.get_config to load YAML files as attribute-style dictionaries accessible via dot notation.
  • Pass the configuration stem to the training script using the --model CLI argument.
  • Save finalized configurations using standard file copy operations after hyper-parameter tuning.

Frequently Asked Questions

What file naming convention should I follow for custom model configurations?

Name your configuration files using lowercase alphanumeric characters and underscores, then place them in configs/models/ with the .yaml extension. When loading via get_config or the CLI --model flag, reference only the stem (e.g., my_custom for configs/models/my_custom.yaml).

Can I nest configuration parameters in the YAML file?

Yes. The get_config function in doc2graph/utils.py recursively converts nested dictionaries into AttrDict instances (lines 57-65). This allows you to organize parameters hierarchically (e.g., optimizer.lr) while still accessing them via attribute syntax in Python.

How does doc2graph handle missing configuration keys?

The model constructors and training pipeline expect specific keys to be present based on the architecture type. If you omit required parameters (such as hidden_dim for GCN models), the code will raise an AttributeError or KeyError when attempting to access cfg.hidden_dim. Always base your custom configurations on existing templates like base.yaml to ensure all required fields are present.

Where is the configuration loaded during the training pipeline?

The loading occurs in doc2graph/main.py when the training script starts. The CLI argument --model passes the configuration name to doc2graph.utils.get_config, which reads the corresponding YAML file from configs/models/ and returns an AttrDict used throughout the training utilities and model initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →