# PyTDC vs DeepChem for ADMET Prediction: Key Differences and When to Use Each

> Compare PyTDC's benchmark datasets and DeepChem's ML stack for ADMET prediction. Learn key differences and choose the right tool for your project.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: comparison
- Published: 2026-05-14

---

**PyTDC provides standardized benchmark datasets and evaluation protocols for comparing ADMET models, while DeepChem offers a complete machine learning stack for building and training custom ADMET predictors from scratch.**

Both PyTDC (Therapeutics Data Commons) and DeepChem are open-source Python toolkits widely used in drug discovery, but they approach ADMET (Absorption, Distribution, Metabolism, Excretion, Toxicity) modeling from fundamentally different architectural perspectives. According to the `scientific-agent-skills` repository, PyTDC excels at standardized benchmarking across multiple endpoints, whereas DeepChem empowers researchers to develop novel neural network architectures for specific ADMET properties.

## Core Architectural Differences

### PyTDC: Benchmark-Centric Standardization

PyTDC treats ADMET prediction as a **benchmarking challenge**. The library does not ship any learning algorithms; instead, it provides a unified API for accessing canonical datasets with predetermined train/validation/test splits. In [`scientific-skills/pytdc/scripts/benchmark_evaluation.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/pytdc/scripts/benchmark_evaluation.py), the `load_benchmark_group()` function returns an `admet_group` object that encapsulates 22 distinct ADMET datasets, each strictly partitioned according to the official TDC leaderboard protocol.

The core philosophy is **model-agnostic evaluation**. You supply predictions from any algorithm—whether a random forest, graph neural network, or custom deep learning model—and PyTDC computes the official metrics via `Evaluator` objects. This design enforces reproducibility by implementing a **5-seed protocol** where splits are fixed per seed, returning mean ± standard deviation across seeds as required for leaderboard submission.

### DeepChem: Model-Centric Development Stack

DeepChem adopts a **full-stack machine learning** approach. As shown in [`scientific-skills/deepchem/scripts/graph_neural_network.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/deepchem/scripts/graph_neural_network.py), the library provides end-to-end components: data loaders (`CSVLoader`), molecular featurizers (`MolGraphConvFeaturizer`), ready-to-use models (GCN, GAT, AttentiveFP, MPNN, D-MPNN), and training loops. Unlike PyTDC, DeepChem expects you to construct the model architecture, define hyperparameters, and manage the training process via `model.fit()` and `model.evaluate()`.

This architecture prioritizes **flexibility over standardization**. Users control every aspect of the pipeline, from featurization to split strategies (e.g., scaffold splitting), making DeepChem ideal for iterative model development rather than leaderboard comparison.

## Data Handling and Splits

PyTDC abstracts data access through **benchmark groups**. When you call `admet_group(path='data/')`, the library automatically downloads ADMET datasets and returns pandas-style objects with splits already separated. The `single_dataset_evaluation()` and `multiple_datasets_evaluation()` functions in the benchmark script handle the tedious work of iterating across seeds and aggregating results.

DeepChem requires explicit data pipeline construction. The `train_on_custom_data()` function in [`graph_neural_network.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/graph_neural_network.py) demonstrates loading data from CSV files or MoleculeNet, applying `MolGraphConvFeaturizer` to convert SMILES strings into graph tensors, and using `ScaffoldSplitter` to create train/valid/test partitions. While this offers granular control over data preparation, it places the burden of split consistency entirely on the user.

## Modeling Approaches

The divergence in modeling philosophy is stark:

- **PyTDC** expects you to generate predictions externally. You pass NumPy arrays of predictions to `group.evaluate(predictions)`, and the library validates your model against the benchmark using canonical metrics (MAE for regression, ROC-AUC for classification).

- **DeepChem** provides parameterized model constructors. The `create_model` utility instantiates `GCNModel`, `AttentiveFPModel`, or `MPNNModel` objects with sensible defaults, then executes training loops. You define the task type (regression or classification), number of epochs, and batch size, iterating until convergence.

## Evaluation Protocols

PyTDC enforces **strict reproducibility standards**. The evaluation logic in [`benchmark_evaluation.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/benchmark_evaluation.py) uses `Evaluator(name='MAE')` to compute metrics, automatically averaging results across five random seeds to match TDC leaderboard requirements. This eliminates variability from random initialization and ensures your results are comparable to published benchmarks.

DeepChem delegates metric selection to the user. While the [`graph_neural_network.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/graph_neural_network.py) script reports standard regression metrics (R², MAE, RMSE) via `dc.metrics`, the specific metrics, split strategies, and random seeds are configurable. This flexibility supports exploratory research but requires careful documentation to ensure reproducibility.

## Code Examples

### Running ADMET Benchmarks with PyTDC

The following example demonstrates the PyTDC workflow for evaluating a model against the official Caco2 permeability benchmark:

```python

# example_admet_tdc.py

from scientific_skills.pytdc.scripts.benchmark_evaluation import (
    load_benchmark_group,
    single_dataset_evaluation,
    multiple_datasets_evaluation,
)

if __name__ == "__main__":
    # Load the ADMET benchmark group

    group = load_benchmark_group()  # Source: scientific-skills/pytdc/scripts/benchmark_evaluation.py

    # Evaluate a single dataset (e.g., Caco2 permeability)

    predictions, results = single_dataset_evaluation(group, dataset_name="Caco2_Wang")
    
    # Evaluate multiple datasets simultaneously

    all_preds, all_res = multiple_datasets_evaluation(group)

```

Under the hood, `load_benchmark_group()` initializes the `admet_group` with fixed splits. The `single_dataset_evaluation()` function iterates through seeds 1-5, computes MAE using `Evaluator`, and returns mean ± standard deviation.

### Training Custom GNNs with DeepChem

This example shows how DeepChem handles end-to-end model training for an ADMET-like regression task:

```python

# example_gnn_deepchem.py

import argparse
from scientific_skills.deepchem.scripts.graph_neural_network import (
    train_on_custom_data,
)

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--data", required=True, help="Path to CSV with SMILES+target")
    parser.add_argument("--targets", nargs="+", default=["solubility"],
                        help="Column name(s) for the ADMET property")
    args = parser.parse_args()

    # Train a GCN on custom ADMET data

    model, test_set = train_on_custom_data(
        data_path=args.data,
        model_type="gcn",          # Options: gcn, gat, attentivefp, mpnn, dmpnn

        task_type="regression",    # ADMET endpoints typically use regression

        target_cols=args.targets,
        smiles_col="smiles",
        n_epochs=50,
    )

```

The `train_on_custom_data()` function orchestrates the full pipeline: `MolGraphConvFeaturizer` generates graph features, `ScaffoldSplitter` partitions the data, and `GCNModel` handles training and evaluation.

### Side-by-Side Comparison

This script illustrates the fundamental workflow difference:

```python

# compare_tdc_vs_deepchem.py

import numpy as np
from scientific_skills.pytdc.scripts.benchmark_evaluation import load_benchmark_group, single_dataset_evaluation
from scientific_skills.deepchem.scripts.graph_neural_network import train_on_custom_data

# PyTDC approach: Benchmark evaluation

group = load_benchmark_group()
_, tdc_results = single_dataset_evaluation(group, dataset_name="Caco2_Wang")
print("PyTDC MAE (mean ± std):", tdc_results["Caco2_Wang"])

# DeepChem approach: End-to-end training

# Requires exporting TDC data to CSV first

model, _ = train_on_custom_data(
    data_path="caco2_export.csv",
    model_type="gcn",
    task_type="regression",
    target_cols=["Y"],
    smiles_col="smiles",
    n_epochs=30,
)

```

PyTDC requires you to generate predictions separately and supply them for evaluation, while DeepChem manages the entire training process.

## Summary

- **PyTDC** is **model-agnostic** and **benchmark-focused**, providing standardized datasets with fixed splits and enforcing a 5-seed evaluation protocol for reproducible leaderboard results.
- **DeepChem** is **model-centric**, offering a complete ML stack with built-in GNN architectures, featurizers, and training loops for custom ADMET model development.
- **PyTDC** handles data splits automatically via `admet_group`, while **DeepChem** requires manual pipeline construction using `CSVLoader` and splitters like `ScaffoldSplitter`.
- **PyTDC** computes canonical metrics (MAE, ROC-AUC) via `Evaluator` objects, whereas **DeepChem** allows flexible metric selection from `dc.metrics`.
- Choose **PyTDC** when comparing models against community standards; choose **DeepChem** when building production ADMET predictors requiring architectural customization.

## Frequently Asked Questions

### Can I use DeepChem models with PyTDC benchmarks?

Yes. The standard workflow involves training your model using DeepChem's infrastructure, generating predictions on PyTDC test sets, and then passing those predictions to PyTDC's `group.evaluate()` function. This hybrid approach leverages DeepChem's model-building capabilities while maintaining the standardized evaluation protocol required for TDC leaderboard submission.

### Why does PyTDC use a 5-seed protocol while DeepChem uses single-seed evaluation?

PyTDC's 5-seed protocol eliminates variance from random train/test splits and model initialization, ensuring that reported performance reflects true algorithmic capability rather than data partitioning luck. This is essential for fair comparison across research groups. DeepChem defaults to single-seed evaluation for faster iteration during development, though you can manually implement multi-seed validation by setting `np.random.seed()` and averaging results across runs.

### Which toolkit is better for production ADMET screening?

**DeepChem** is generally preferred for production pipelines because it provides trained model objects that can be serialized and deployed for inference on new chemical libraries. PyTDC is primarily designed for research benchmarking and does not export trained models; it evaluates pre-computed predictions against fixed test sets.

### Do both libraries support graph neural networks for ADMET prediction?

Only **DeepChem** natively supports GNNs with implementations like GCN, GAT, AttentiveFP, and MPNN in [`scientific-skills/deepchem/scripts/graph_neural_network.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/deepchem/scripts/graph_neural_network.py). PyTDC is featurization- and model-agnostic, meaning you can use GNNs to generate predictions, but the library itself does not provide neural network implementations. You must train GNNs using PyTorch Geometric, DeepChem, or another framework, then feed the predictions into PyTDC's evaluation framework.