How to Load ESMProtein from PDB Files Using from_pdb

Use ESMProtein.from_pdb() to parse PDB files into structured protein objects with sequence, coordinates, and confidence scores in a single line of code.

The ESMProtein class serves as the central data structure in the Biohub/esm repository, unifying protein sequences, structural coordinates, and metadata for downstream machine learning workflows. When you need to load ESMProtein from PDB files, the SDK provides a dedicated factory method that handles parsing, chain selection, and data normalization automatically.

Understanding the ESMProtein.from_pdb Method

In esm/sdk/api.py, the ESMProtein dataclass exposes the from_pdb class method between lines 68-95. This function acts as a high-level interface that delegates to lower-level parsers while standardizing the output format.

The method signature accepts:

  • pdb_path: A string path or file-like object pointing to the PDB file
  • chain_id: Either "all" for the entire complex, "detect" for automatic first-chain selection, or a specific author-assigned chain identifier
  • is_predicted: A boolean flag indicating whether the structure contains B-factor confidence values typical of AlphaFold outputs

Internally, from_pdb routes to ProteinComplex.from_pdb or ProteinChain.from_pdb (located in esm/utils/structure/protein_complex.py and esm/utils/structure/protein_chain.py respectively), then converts the intermediate objects using ESMProtein.from_protein_complex or ESMProtein.from_protein_chain.

Loading PDB Files: Step-by-Step Implementation

To begin, import the class from the SDK entry point defined at lines 24-28 in esm/sdk/api.py:

from esm.sdk.api import ESMProtein

Loading All Chains

Pass chain_id="all" to import the complete protein complex:

pdb_path = "data/1abc.pdb"
protein = ESMProtein.from_pdb(pdb_path, chain_id="all")

print(f"Sequence length: {len(protein)}")
print(f"Number of atoms: {protein.coordinates.shape[0]}")

This returns an ESMProtein instance where coordinates include all chains present in the asymmetric unit.

Loading a Specific Chain

Specify an author-assigned chain ID to isolate a single polypeptide:

protein_A = ESMProtein.from_pdb(pdb_path, chain_id="A")
print(protein_A.sequence[:30])

Chain Selection Strategies

The chain_id parameter supports three distinct modes of operation:

  • "all": Loads the entire ProteinComplex, preserving inter-chain interactions and stoichiometry
  • "detect": Automatically selects the first chain encountered in the PDB file
  • Specific ID: Any string matching the PDB author chain identifier (e.g., "A", "B", "H" for heavy chains)

This flexibility allows you to process multimeric assemblies or isolate specific subunits without preprocessing the PDB file externally.

Handling Predicted Structures (AlphaFold)

When working with computational models that include confidence metrics, set is_predicted=True:

protein_pred = ESMProtein.from_pdb(
    pdb_path, 
    chain_id="all", 
    is_predicted=True
)
print("pLDDT tensor shape:", protein_pred.plddt.shape)

This flag ensures the parser correctly interprets B-factor columns as pLDDT (predicted Local Distance Difference Test) scores rather than crystallographic temperature factors, populating the .plddt attribute with per-residue confidence values.

Working with the Result

The returned ESMProtein object exposes standardized attributes ready for ESM model consumption:

  • sequence: The amino acid sequence as a string
  • coordinates: Atom coordinates as a tensor with shape (num_atoms, 3)
  • plddt: Per-residue confidence scores (when is_predicted=True)

You can feed these directly into tokenization pipelines:

from esm.tokenization import get_esm3_model_tokenizers

tokenizers = get_esm3_model_tokenizers()
seq_tokens = tokenizers.sequence_tokenizer.encode(protein.sequence)

print("First 10 tokens:", seq_tokens[:10])

This integration bridges structural biology data and the ESM-3 model architecture without manual data conversion.

Summary

  • ESMProtein.from_pdb() in esm/sdk/api.py provides the canonical entry point to load ESMProtein from PDB files
  • The method supports both experimental and predicted structures via the is_predicted parameter
  • Chain selection options include "all", "detect", or specific author IDs for flexible multimer handling
  • Output objects contain .sequence, .coordinates, and .plddt attributes compatible with downstream ESM tokenizers and inference pipelines
  • Internal parsing delegates to ProteinComplex and ProteinChain utilities in esm/utils/structure/

Frequently Asked Questions

What is the difference between chain_id="all" and chain_id="detect"?

When you specify chain_id="all", the method constructs a ProteinComplex containing every chain present in the PDB file, preserving the biological assembly. Using chain_id="detect" extracts only the first chain encountered, which is useful for monomeric structures or when the specific chain ID is unknown. Both paths ultimately convert to ESMProtein format, but "all" retains inter-chain spatial relationships.

How do I load AlphaFold predicted structures with confidence scores?

Set the is_predicted=True parameter when calling ESMProtein.from_pdb(). This instructs the parser in esm/utils/structure/protein_chain.py to interpret B-factor values as pLDDT confidence scores rather than crystallographic B-factors. The resulting object will populate the .plddt attribute with a tensor of per-residue confidence values ranging from 0-100.

Can I load PDB content from a file-like object instead of a file path?

Yes. The from_pdb method accepts any file-like object that implements a read interface, not just string file paths. You can pass an opened file handle (open("structure.pdb", "r")), a BytesIO object, or any compatible stream. This is particularly useful when fetching structures from databases without writing intermediate files to disk.

Which source files contain the underlying implementation?

The public API resides in esm/sdk/api.py (lines 68-95), which orchestrates the parsing logic. The actual PDB parsing implementations are found in esm/utils/structure/protein_complex.py for multimeric structures and esm/utils/structure/protein_chain.py for single chains. These modules handle coordinate extraction, residue mapping, and confidence score parsing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →