# How to Load ESMProtein from PDB Files Using from_pdb

> Easily load ESMProtein from PDB files using ESMProtein.from_pdb() to get sequence coordinates and confidence scores in one line.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**Use `ESMProtein.from_pdb()` to parse PDB files into structured protein objects with sequence, coordinates, and confidence scores in a single line of code.**

The `ESMProtein` class serves as the central data structure in the **Biohub/esm** repository, unifying protein sequences, structural coordinates, and metadata for downstream machine learning workflows. When you need to **load ESMProtein from PDB** files, the SDK provides a dedicated factory method that handles parsing, chain selection, and data normalization automatically.

## Understanding the ESMProtein.from_pdb Method

In [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py), the `ESMProtein` dataclass exposes the `from_pdb` class method between lines 68-95. This function acts as a high-level interface that delegates to lower-level parsers while standardizing the output format.

The method signature accepts:

- **pdb_path**: A string path or file-like object pointing to the PDB file
- **chain_id**: Either `"all"` for the entire complex, `"detect"` for automatic first-chain selection, or a specific author-assigned chain identifier
- **is_predicted**: A boolean flag indicating whether the structure contains B-factor confidence values typical of AlphaFold outputs

Internally, `from_pdb` routes to `ProteinComplex.from_pdb` or `ProteinChain.from_pdb` (located in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) and [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) respectively), then converts the intermediate objects using `ESMProtein.from_protein_complex` or `ESMProtein.from_protein_chain`.

## Loading PDB Files: Step-by-Step Implementation

To begin, import the class from the SDK entry point defined at lines 24-28 in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py):

```python
from esm.sdk.api import ESMProtein

```

### Loading All Chains

Pass `chain_id="all"` to import the complete protein complex:

```python
pdb_path = "data/1abc.pdb"
protein = ESMProtein.from_pdb(pdb_path, chain_id="all")

print(f"Sequence length: {len(protein)}")
print(f"Number of atoms: {protein.coordinates.shape[0]}")

```

This returns an `ESMProtein` instance where coordinates include all chains present in the asymmetric unit.

### Loading a Specific Chain

Specify an author-assigned chain ID to isolate a single polypeptide:

```python
protein_A = ESMProtein.from_pdb(pdb_path, chain_id="A")
print(protein_A.sequence[:30])

```

## Chain Selection Strategies

The `chain_id` parameter supports three distinct modes of operation:

- **`"all"`**: Loads the entire `ProteinComplex`, preserving inter-chain interactions and stoichiometry
- **`"detect"`**: Automatically selects the first chain encountered in the PDB file
- **Specific ID**: Any string matching the PDB author chain identifier (e.g., `"A"`, `"B"`, `"H"` for heavy chains)

This flexibility allows you to process multimeric assemblies or isolate specific subunits without preprocessing the PDB file externally.

## Handling Predicted Structures (AlphaFold)

When working with computational models that include confidence metrics, set `is_predicted=True`:

```python
protein_pred = ESMProtein.from_pdb(
    pdb_path, 
    chain_id="all", 
    is_predicted=True
)
print("pLDDT tensor shape:", protein_pred.plddt.shape)

```

This flag ensures the parser correctly interprets B-factor columns as **pLDDT** (predicted Local Distance Difference Test) scores rather than crystallographic temperature factors, populating the `.plddt` attribute with per-residue confidence values.

## Working with the Result

The returned `ESMProtein` object exposes standardized attributes ready for ESM model consumption:

- **sequence**: The amino acid sequence as a string
- **coordinates**: Atom coordinates as a tensor with shape `(num_atoms, 3)`
- **plddt**: Per-residue confidence scores (when `is_predicted=True`)

You can feed these directly into tokenization pipelines:

```python
from esm.tokenization import get_esm3_model_tokenizers

tokenizers = get_esm3_model_tokenizers()
seq_tokens = tokenizers.sequence_tokenizer.encode(protein.sequence)

print("First 10 tokens:", seq_tokens[:10])

```

This integration bridges structural biology data and the ESM-3 model architecture without manual data conversion.

## Summary

- **`ESMProtein.from_pdb()`** in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) provides the canonical entry point to **load ESMProtein from PDB** files
- The method supports both experimental and predicted structures via the `is_predicted` parameter
- Chain selection options include `"all"`, `"detect"`, or specific author IDs for flexible multimer handling
- Output objects contain `.sequence`, `.coordinates`, and `.plddt` attributes compatible with downstream ESM tokenizers and inference pipelines
- Internal parsing delegates to `ProteinComplex` and `ProteinChain` utilities in `esm/utils/structure/`

## Frequently Asked Questions

### What is the difference between chain_id="all" and chain_id="detect"?

When you specify `chain_id="all"`, the method constructs a `ProteinComplex` containing every chain present in the PDB file, preserving the biological assembly. Using `chain_id="detect"` extracts only the first chain encountered, which is useful for monomeric structures or when the specific chain ID is unknown. Both paths ultimately convert to `ESMProtein` format, but `"all"` retains inter-chain spatial relationships.

### How do I load AlphaFold predicted structures with confidence scores?

Set the `is_predicted=True` parameter when calling `ESMProtein.from_pdb()`. This instructs the parser in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) to interpret B-factor values as pLDDT confidence scores rather than crystallographic B-factors. The resulting object will populate the `.plddt` attribute with a tensor of per-residue confidence values ranging from 0-100.

### Can I load PDB content from a file-like object instead of a file path?

Yes. The `from_pdb` method accepts any file-like object that implements a read interface, not just string file paths. You can pass an opened file handle (`open("structure.pdb", "r")`), a `BytesIO` object, or any compatible stream. This is particularly useful when fetching structures from databases without writing intermediate files to disk.

### Which source files contain the underlying implementation?

The public API resides in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) (lines 68-95), which orchestrates the parsing logic. The actual PDB parsing implementations are found in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) for multimeric structures and [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) for single chains. These modules handle coordinate extraction, residue mapping, and confidence score parsing.