How to Handle Multi-Chain Protein Complexes with Chain IDs in ESMProtein

To handle multi-chain protein complexes in ESMProtein, create individual ProteinChain objects with explicit chain_id parameters, assemble them into a ProteinComplex, and convert using ESMProtein.from_protein_complex() to preserve chain identifiers throughout tokenization, model inference, and PDB export.

The ESM (Evolutionary Scale Modeling) repository from Biohub provides the ESMProtein class as the high-level data structure for feeding protein data into inference APIs. When working with multi-chain protein complexes—such as antibody-antigen pairs or multi-subunit enzymes—maintaining chain identifiers ensures accurate structural prediction and proper downstream analysis. The following workflow uses the actual implementation from the Biohub/esm source code to keep chain IDs intact from input through to PDB export.

Step 1: Create ProteinChain Objects with Chain IDs

Individual biological chains must first be instantiated as ProteinChain objects with explicit chain identifiers. The ProteinChain class stores the amino-acid sequence and a mutable chain_id attribute according to the implementation in esm/utils/structure/protein_chain.py (lines 148-158).

Use the from_sequence() class method and specify the chain_id parameter for each chain:

from esm.utils.structure.protein_chain import ProteinChain

# Create chains with author-provided IDs (e.g., "A", "B")

chain_A = ProteinChain.from_sequence(
    sequence="MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHF",
    chain_id="A",
)

chain_B = ProteinChain.from_sequence(
    sequence="GHHHHHHSSGVDLGTENLYFQSMVSKGEEDNMASL",
    chain_id="B",
)

Step 2: Assemble Chains into a ProteinComplex

Combine the individual chains into a ProteinComplex container. This class maintains an ordered list of ProteinChain objects and generates an np.ndarray called chain_id that maps every residue token to its corresponding chain index, as implemented in esm/utils/structure/protein_complex.py (lines 112-124).

from esm.utils.structure.protein_complex import ProteinComplex

complex = ProteinComplex(
    chains=[chain_A, chain_B],
    metadata=None,  # Optional: can store organism, stoichiometry, etc.

)

The ProteinComplex class also provides helper methods for ID management. You can call add_prefix_to_chain_ids() or normalize_chain_ids_for_pdb() if you need to conform to PDB limitations or avoid naming collisions.

Step 3: Convert to ESMProtein Using from_protein_complex()

Pass the assembled complex to ESMProtein.from_protein_complex() to generate the tokenized representation required by the model. This method extracts the sequences, builds the tokenized representation, and attaches a chain_id token array to the internal MolecularComplex object, according to the implementation in esm/sdk/api.py (lines 119-128).

from esm.sdk.api import ESMProtein

esm_protein = ESMProtein.from_protein_complex(complex)

The resulting ESMProtein instance maintains the chain identifier information at the token level, ensuring that multi-chain complexes are processed correctly during inference.

Step 4: Access Chain IDs After Inference

After model generation or decoding, retrieve the per-token chain IDs through the molecular_complex attribute. The chain_id array is stored as an np.int64 index that can be mapped back to the original author letters via the metadata.chain_lookup dictionary inside the ProteinComplex, as found in esm/utils/structure/protein_complex.py (lines 305-313).


# Access token-wise chain IDs (shape: [N_tokens])

chain_id_per_token = esm_protein.molecular_complex.chain_id

# Export to PDB while preserving chain letters

pdb_bytes = esm_protein.to_pdb()

Complete End-to-End Example

The following example demonstrates the full workflow for a three-chain complex:

from esm.utils.structure.protein_chain import ProteinChain
from esm.utils.structure.protein_complex import ProteinComplex
from esm.sdk.api import ESMProtein

# Define three chains with distinct IDs

chains = [
    ProteinChain.from_sequence(
        "MKTIIALSYIFCLVFADYKDDDDK", chain_id="A"),
    ProteinChain.from_sequence(
        "GHHHHHHSSGVDLGTENLYFQSM", chain_id="B"),
    ProteinChain.from_sequence(
        "VTVGETRKVAKVST", chain_id="C"),
]

# Assemble into complex

complex = ProteinComplex(chains=chains)

# Convert to ESMProtein format

esm_prot = ESMProtein.from_protein_complex(complex)

# Inspect chain IDs in the tokenized representation

print("Token-wise chain IDs:", esm_prot.molecular_complex.chain_id)

# Export to PDB string

pdb_output = esm_prot.to_pdb()

Implementation Details and Key Source Files

Understanding the source architecture helps debug multi-chain workflows:

  • esm/sdk/api.py – Contains ESMProtein.from_protein_complex(), the entry point for converting a ProteinComplex into an ESMProtein (lines 119-128).
  • esm/utils/structure/protein_chain.py – Implements ProteinChain with the chain_id attribute and from_sequence() method (lines 148-158).
  • esm/utils/structure/protein_complex.py – Defines ProteinComplex, which manages the chain_id array and the metadata.chain_lookup dictionary for mapping indices to author letters (lines 112-124, 305-313).
  • esm/utils/structure/molecular_complex.py – Houses the internal token-level representation that stores the per-token chain_id array used during model inference.

Summary

  • Use ProteinChain.from_sequence() with the chain_id parameter to explicitly label each biological chain before assembly.
  • Assemble chains into a ProteinComplex to generate the token-to-chain mapping array required by the model.
  • Call ESMProtein.from_protein_complex() to convert complexes while preserving chain identifiers in the internal MolecularComplex representation.
  • Access chain information via esm_protein.molecular_complex.chain_id or export to PDB using to_pdb() to retain original chain letters.

Frequently Asked Questions

What happens if I don't specify chain_id when creating ProteinChain objects?

If you omit the chain_id parameter, ProteinChain may assign default identifiers or leave the attribute unset, which can cause chain tracking to fail during ProteinComplex assembly. Always provide explicit chain_id values (e.g., "A", "B", "C") to ensure the chain_id array in ProteinComplex correctly maps tokens to their source chains.

Can I modify chain IDs after creating a ProteinComplex?

Yes. The ProteinComplex class provides helper methods such as add_prefix_to_chain_ids() and normalize_chain_ids_for_pdb() to modify chain identifiers before conversion to ESMProtein. You can also manually adjust the chain_id attribute on individual ProteinChain objects since it is mutable, then reassemble the complex if needed.

How does the chain_id array relate to actual PDB chain letters?

The chain_id array stored in molecular_complex consists of np.int64 indices representing chain membership for each token. The ProteinComplex maintains a metadata.chain_lookup dictionary that maps these integer indices back to the original author-provided chain letters (e.g., {0: "A", 1: "B"}), which to_pdb() uses to generate standard PDB output with correct chain identifiers.

Is there a limit to the number of chains supported in ESMProtein?

While the ProteinComplex class technically supports an arbitrary number of chains via its Python list structure, practical limits depend on the specific ESM model's maximum sequence length and GPU memory constraints. The chain ID mapping system itself uses standard NumPy arrays and can theoretically handle hundreds of chains, though most inference use cases involve complexes with fewer than 62 chains to comply with single-character PDB chain ID conventions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →