# How to Integrate ESMC Embeddings with Molecular Dynamics Simulations

> Integrate ESM-C embeddings into molecular dynamics simulations. Generate per-residue embeddings, map them to structures, and export for OpenMM or GROMACS restraint weights or collective variables in MD engines.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**You can integrate ESM‑C embeddings with molecular dynamics (MD) simulations by generating per‑residue embeddings with the `ESMC` model, mapping them to structural coordinates via `MolecularComplex`, and exporting the enriched data to mmCIF or PDB format for use as restraint weights or collective variables in MD engines like OpenMM or GROMACS.**

The Biohub/esm repository provides a complete toolkit for bridging protein language models and classical simulation workflows. By combining the transformer‑based **ESM‑C** architecture with the structural utilities in `esm.utils.structure`, you can directly incorporate sequence‑derived embeddings into molecular dynamics pipelines without external alignment tools.

## Overview of the Integration Pipeline

The workflow relies on three core components that handle sequence encoding, structure representation, and data export.

### Core Components

- **`esm.models.esmc.ESMC`** (located in [`esm/models/esmc.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmc.py)): The transformer model that generates per‑residue embeddings. The `logits()` method accepts a `LogitsConfig` with `return_embeddings=True` to output an `ESMCOutput` object containing the embedding tensor.

- **`esm.utils.structure.molecular_complex.MolecularComplex`** (located in [`esm/utils/structure/molecular_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/molecular_complex.py)): A flat‑atom representation that maintains exact token‑to‑atom mapping. The `sequence` attribute preserves the same token order as the embedding tensor, enabling direct assignment of features to residues.

- **`esm.utils.structure.protein_complex.ProteinComplex`** (located in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py)): Provides the classic `atom37` view and `to_pdb()` export method required by traditional MD packages.

## Step‑by‑Step Implementation

### 1. Generate Per‑Residue Embeddings

Load a pretrained ESM‑C model and encode your target sequence. Request embeddings via the `LogitsConfig` API.

```python
import torch
from esm.models.esmc import ESMC
from esm.sdk.api import LogitsConfig, ESMProtein

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = ESMC.from_pretrained("esmc_600m", device=device)

protein = ESMProtein(sequence="MKTLLILAVLLVVLGVAGSS")
protein_tensor = model.encode(protein)

logits_out = model.logits(
    protein_tensor,
    LogitsConfig(sequence=True, return_embeddings=True)
)

embeddings = logits_out.embeddings  # Shape: (L, d_model)

```

### 2. Load the Corresponding Structure

Import your PDB or mmCIF file using `MolecularComplex.from_mmcif()`. The token order in `MolecularComplex.sequence` exactly matches the embedding order returned by the model.

```python
from esm.utils.structure.molecular_complex import MolecularComplex

mol_complex = MolecularComplex.from_mmcif("data/example.pdb")
assert len(mol_complex.sequence) == embeddings.shape[0]  # Alignment verified

```

### 3. Attach Embeddings and Export

Store scalar values derived from the embeddings (such as the L2 norm) in the `plddt` field, which maps to the B‑factor column during export. This allows MD engines to read per‑residue confidence or restraint weights.

```python
import numpy as np

# Calculate per‑residue embedding norm

b_factor = embeddings.norm(dim=1).cpu().numpy()

# Broadcast to all atoms belonging to each token

atom_bf = []
for start, end in mol_complex.token_to_atoms:
    atom_bf.extend([b_factor[i]] * (end - start))

mol_complex.plddt = np.array(atom_bf, dtype=np.float32)

# Export to mmCIF for MD engines

mmcif_str = mol_complex.to_mmcif()
with open("complex_with_emb.mmcif", "w") as f:
    f.write(mmcif_str)

```

Alternatively, convert to `ProteinComplex` for PDB export:

```python
protein_complex = mol_complex.to_protein_complex()
pdb_str = protein_complex.to_pdb()

```

## Practical MD Applications

### Using Embeddings as Harmonic Restraints (OpenMM)

You can inject embedding‑derived values into a `CustomExternalForce` to create position restraints that scale with the sequence embedding magnitude.

```python
from openmm import *
from openmm.app import PDBFile, ForceField

pdb = PDBFile("complex_with_emb.pdb")
forcefield = ForceField('amber14-all.xml')
system = forcefield.createSystem(pdb.topology)

# Create restraint using B‑factor values as force constants

restraint = CustomExternalForce('0.5*k*periodicdistance(x,y,z,x0,y0,z0)^2')
restraint.addPerParticleParameter('k')
restraint.addPerParticleParameter('x0')
restraint.addPerParticleParameter('y0')
restraint.addPerParticleParameter('z0')

# Assume b_factors loaded from the PDB file

for i, atom in enumerate(pdb.topology.atoms()):
    k = b_factors[i] * 0.1  # Scale to kJ·mol⁻¹·nm⁻²

    pos = pdb.positions[i]
    restraint.addParticle(i, [k, pos.x, pos.y, pos.z])

system.addForce(restraint)

```

### Coarse‑Grained Simulations with Per‑Token Coordinates

For residue‑level coarse‑grained MD, compute centroids from the flat atom arrays and write an XYZ file.

```python
import numpy as np

centroids = []
for start, end in mol_complex.token_to_atoms:
    token_coords = mol_complex.atom_positions[start:end]
    centroids.append(token_coords.mean(axis=0))

centroids = np.stack(centroids)

with open("cg_structure.xyz", "w") as f:
    f.write(f"{len(centroids)}\n")
    f.write("CG model from ESM‑C embeddings\n")
    for i, (x, y, z) in enumerate(centroids):
        f.write(f"C{i+1:04d} {x:.3f} {y:.3f} {z:.3f}\n")

```

## Key Source Files Reference

| File | Relevance to MD Integration |
|------|------------------------------|
| [`esm/models/esmc.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmc.py) | Implements `ESMC.from_pretrained()` and the `logits()` API used to extract embeddings. |
| [`esm/utils/structure/molecular_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/molecular_complex.py) | Contains `from_mmcif()`, `to_mmcif()`, and the `token_to_atoms` mapping critical for alignment. |
| [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) | Provides `to_pdb()` for exporting to legacy MD formats. |
| [`esm/tokenization/sequence_tokenizer.py`](https://github.com/Biohub/esm/blob/main/esm/tokenization/sequence_tokenizer.py) | Supplies `EsmSequenceTokenizer` for converting sequences to model inputs. |

## Summary

- **ESM‑C** generates per‑residue embeddings via `model.logits(..., LogitsConfig(return_embeddings=True))`.
- **`MolecularComplex`** maintains token‑atom alignment automatically, eliminating the need for manual sequence‑structure alignment.
- Store embedding‑derived scalars in the **`plddt`** field to repurpose the B‑factor column for restraint weights.
- Export to **mmCIF** or **PDB** using `to_mmcif()` or `to_pdb()` for direct consumption by OpenMM, GROMACS, or AMBER.

## Frequently Asked Questions

### How do I ensure token‑to‑residue alignment when integrating ESMC embeddings with MD?

The `MolecularComplex.sequence` attribute uses the exact same tokenization as the `ESMC` model. When you load a structure with `MolecularComplex.from_mmcif()`, the resulting object preserves the token order, allowing direct indexing between `embeddings[i]` and the residue at `mol_complex.sequence[i]`.

### Can I use ESMC embeddings directly as force field parameters?

Yes, you can derive scalar values from embeddings—such as the L2 norm, variance, or specific dimension values—and map them to per‑atom force constants or collective variable biases. The Biohub/esm source code demonstrates this by storing derived values in the `plddt` field, which exports to the B‑factor column readable by most MD engines.

### What file format should I use to export structures for GROMACS or AMBER?

Use `MolecularComplex.to_mmcif()` for GROMACS, as it supports modern mmCIF format with explicit chain IDs. For AMBER or legacy workflows, convert to `ProteinComplex` via `to_protein_complex()` and call `to_pdb()` to generate standard PDB files that preserve the embedding‑derived B‑factors.

### Is GPU acceleration required for generating embeddings before MD simulations?

While GPU acceleration significantly speeds up embedding generation via `ESMC`, it is not strictly required. The model supports CPU inference, and once the embeddings are exported to a structure file, the subsequent MD simulation runs independently on any hardware supported by your chosen MD engine.