# How to Handle Multi-Chain Protein Complexes with Chain IDs in ESMProtein

> Learn to handle multi-chain protein complexes in ESMProtein. Create ProteinChain objects with chain IDs, assemble them into a ProteinComplex, and convert to maintain identifiers for analysis and export.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**To handle multi-chain protein complexes in ESMProtein, create individual `ProteinChain` objects with explicit `chain_id` parameters, assemble them into a `ProteinComplex`, and convert using `ESMProtein.from_protein_complex()` to preserve chain identifiers throughout tokenization, model inference, and PDB export.**

The ESM (Evolutionary Scale Modeling) repository from Biohub provides the `ESMProtein` class as the high-level data structure for feeding protein data into inference APIs. When working with **multi-chain protein complexes**—such as antibody-antigen pairs or multi-subunit enzymes—maintaining chain identifiers ensures accurate structural prediction and proper downstream analysis. The following workflow uses the actual implementation from the Biohub/esm source code to keep chain IDs intact from input through to PDB export.

## Step 1: Create ProteinChain Objects with Chain IDs

Individual biological chains must first be instantiated as `ProteinChain` objects with explicit chain identifiers. The `ProteinChain` class stores the amino-acid sequence and a mutable `chain_id` attribute according to the implementation in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) (lines 148-158).

Use the `from_sequence()` class method and specify the `chain_id` parameter for each chain:

```python
from esm.utils.structure.protein_chain import ProteinChain

# Create chains with author-provided IDs (e.g., "A", "B")

chain_A = ProteinChain.from_sequence(
    sequence="MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHF",
    chain_id="A",
)

chain_B = ProteinChain.from_sequence(
    sequence="GHHHHHHSSGVDLGTENLYFQSMVSKGEEDNMASL",
    chain_id="B",
)

```

## Step 2: Assemble Chains into a ProteinComplex

Combine the individual chains into a `ProteinComplex` container. This class maintains an ordered list of `ProteinChain` objects and generates an `np.ndarray` called `chain_id` that maps every residue token to its corresponding chain index, as implemented in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) (lines 112-124).

```python
from esm.utils.structure.protein_complex import ProteinComplex

complex = ProteinComplex(
    chains=[chain_A, chain_B],
    metadata=None,  # Optional: can store organism, stoichiometry, etc.

)

```

The `ProteinComplex` class also provides helper methods for ID management. You can call `add_prefix_to_chain_ids()` or `normalize_chain_ids_for_pdb()` if you need to conform to PDB limitations or avoid naming collisions.

## Step 3: Convert to ESMProtein Using from_protein_complex()

Pass the assembled complex to `ESMProtein.from_protein_complex()` to generate the tokenized representation required by the model. This method extracts the sequences, builds the tokenized representation, and attaches a `chain_id` token array to the internal `MolecularComplex` object, according to the implementation in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) (lines 119-128).

```python
from esm.sdk.api import ESMProtein

esm_protein = ESMProtein.from_protein_complex(complex)

```

The resulting `ESMProtein` instance maintains the chain identifier information at the token level, ensuring that multi-chain complexes are processed correctly during inference.

## Step 4: Access Chain IDs After Inference

After model generation or decoding, retrieve the per-token chain IDs through the `molecular_complex` attribute. The `chain_id` array is stored as an `np.int64` index that can be mapped back to the original author letters via the `metadata.chain_lookup` dictionary inside the `ProteinComplex`, as found in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) (lines 305-313).

```python

# Access token-wise chain IDs (shape: [N_tokens])

chain_id_per_token = esm_protein.molecular_complex.chain_id

# Export to PDB while preserving chain letters

pdb_bytes = esm_protein.to_pdb()

```

## Complete End-to-End Example

The following example demonstrates the full workflow for a three-chain complex:

```python
from esm.utils.structure.protein_chain import ProteinChain
from esm.utils.structure.protein_complex import ProteinComplex
from esm.sdk.api import ESMProtein

# Define three chains with distinct IDs

chains = [
    ProteinChain.from_sequence(
        "MKTIIALSYIFCLVFADYKDDDDK", chain_id="A"),
    ProteinChain.from_sequence(
        "GHHHHHHSSGVDLGTENLYFQSM", chain_id="B"),
    ProteinChain.from_sequence(
        "VTVGETRKVAKVST", chain_id="C"),
]

# Assemble into complex

complex = ProteinComplex(chains=chains)

# Convert to ESMProtein format

esm_prot = ESMProtein.from_protein_complex(complex)

# Inspect chain IDs in the tokenized representation

print("Token-wise chain IDs:", esm_prot.molecular_complex.chain_id)

# Export to PDB string

pdb_output = esm_prot.to_pdb()

```

## Implementation Details and Key Source Files

Understanding the source architecture helps debug multi-chain workflows:

- **[`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py)** – Contains `ESMProtein.from_protein_complex()`, the entry point for converting a `ProteinComplex` into an `ESMProtein` (lines 119-128).
- **[`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py)** – Implements `ProteinChain` with the `chain_id` attribute and `from_sequence()` method (lines 148-158).
- **[`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py)** – Defines `ProteinComplex`, which manages the `chain_id` array and the `metadata.chain_lookup` dictionary for mapping indices to author letters (lines 112-124, 305-313).
- **[`esm/utils/structure/molecular_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/molecular_complex.py)** – Houses the internal token-level representation that stores the per-token `chain_id` array used during model inference.

## Summary

- **Use `ProteinChain.from_sequence()`** with the `chain_id` parameter to explicitly label each biological chain before assembly.
- **Assemble chains** into a `ProteinComplex` to generate the token-to-chain mapping array required by the model.
- **Call `ESMProtein.from_protein_complex()`** to convert complexes while preserving chain identifiers in the internal `MolecularComplex` representation.
- **Access chain information** via `esm_protein.molecular_complex.chain_id` or export to PDB using `to_pdb()` to retain original chain letters.

## Frequently Asked Questions

### What happens if I don't specify chain_id when creating ProteinChain objects?

If you omit the `chain_id` parameter, `ProteinChain` may assign default identifiers or leave the attribute unset, which can cause chain tracking to fail during `ProteinComplex` assembly. Always provide explicit `chain_id` values (e.g., "A", "B", "C") to ensure the `chain_id` array in `ProteinComplex` correctly maps tokens to their source chains.

### Can I modify chain IDs after creating a ProteinComplex?

Yes. The `ProteinComplex` class provides helper methods such as `add_prefix_to_chain_ids()` and `normalize_chain_ids_for_pdb()` to modify chain identifiers before conversion to `ESMProtein`. You can also manually adjust the `chain_id` attribute on individual `ProteinChain` objects since it is mutable, then reassemble the complex if needed.

### How does the chain_id array relate to actual PDB chain letters?

The `chain_id` array stored in `molecular_complex` consists of `np.int64` indices representing chain membership for each token. The `ProteinComplex` maintains a `metadata.chain_lookup` dictionary that maps these integer indices back to the original author-provided chain letters (e.g., {0: "A", 1: "B"}), which `to_pdb()` uses to generate standard PDB output with correct chain identifiers.

### Is there a limit to the number of chains supported in ESMProtein?

While the `ProteinComplex` class technically supports an arbitrary number of chains via its Python list structure, practical limits depend on the specific ESM model's maximum sequence length and GPU memory constraints. The chain ID mapping system itself uses standard NumPy arrays and can theoretically handle hundreds of chains, though most inference use cases involve complexes with fewer than 62 chains to comply with single-character PDB chain ID conventions.