# How ESMFold2 Handles DNA and RNA Sequences with Protein Chains

> Discover how ESMFold2 predicts mixed protein-DNA-RNA complexes. Learn about its type-aware input pipeline for unified token representations and end-to-end structure prediction of heterogeneous assemblies.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: deep-dive
- Published: 2026-05-30

---

**ESMFold2 predicts structures of mixed protein-DNA-RNA complexes by converting nucleic acid sequences into unified token representations through a type-aware input pipeline that assigns entity IDs, performs nucleotide-specific tokenization, and constructs chemistry-aware reference frames, enabling end-to-end structure prediction of heterogeneous biomolecular assemblies.**

ESMFold2 extends beyond protein-only folding to handle arbitrary combinations of protein, DNA, and RNA chains within a single structure prediction run. According to the Biohub/esm source code, the model achieves this capability through a sophisticated input preparation pipeline implemented primarily in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py), which treats nucleic acids as first-class citizens alongside protein polymers.

## Input Normalization and Chain Construction

The process begins in [`esm/models/esmfold2/processor.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/processor.py) where the `ESMFold2InputBuilder` class orchestrates the preprocessing. The `clean_esmfold2_input` function first expands chain-break tokens (`|`) into separate input objects and validates that covalent-bond specifications include explicit chain IDs. This ensures that mixed complexes are correctly partitioned before tokenization.

Next, `prepare_esmfold2_input` calls `build_chains_from_input` to iterate over `StructurePredictionInput.sequences`. This function detects the concrete type of each entry—`ProteinInput`, `DNAInput`, or `RNAInput`—and assigns an **entity id**, **asym id**, and **sym id** to each chain. Entity deduplication ensures identical sequences share the same `entity_id`, while distinct `asym_id` values maintain separate chain instances.

## Nucleotide Tokenization Strategy

In [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py), the `tokenize_nucleotide` function handles the molecular specifics of nucleic acids. The function converts single-letter nucleotide codes to three-letter residue names (e.g., `A → ADE`, `C → CYT`) and creates **one token per standard residue** containing all heavy atoms defined in `DNA_HEAVY_ATOMS` or `RNA_HEAVY_ATOMS` from [`esm/models/esmfold2/constants.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/constants.py). Modified or non-standard nucleotides fall back to atom-level tokenization, generating one token per atom.

## Frame Construction and Representative Atoms

After tokenization, `compute_frame_indices` establishes local reference frames for each token. For nucleic acid residues, these frames utilize the sugar ring backbone atoms **C1′, C3′, and C4′** (or their RNA analogs), distinct from the N-Cα-C frames used for proteins.

For distance-based structure generation, `compute_representative_atoms` selects specific atoms per nucleotide for distogram conditioning:
- **Purines (A, G)** → **C4** (with C1′ fallback)
- **Pyrimidines (C, T, U)** → **C2** (with C1′ fallback)
- **Unknown nucleotides** → **C1′**

## MSA Handling for Nucleic Acids

As implemented in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py), only protein chains generate true multiple-sequence alignments (MSAs). DNA and RNA chains receive **gap-only MSAs** where the first row contains the target residue type and subsequent rows are filled with `MSA_GAP_TOKEN_ID`. This satisfies the diffusion model's expectation of fixed-size tensors while bypassing evolutionary feature extraction for nucleic acids.

## Tensor Assembly for Mixed Complexes

The final stage packs all per-token and per-atom data into unified tensors including `coords`, `ref_element`, `ref_atom_name_chars`, `token_bonds`, and `distogram_atom_idx`. These tensors contain mixed molecular types—protein, DNA, RNA, or ligands—which the diffusion model processes uniformly regardless of origin. The `ChainInfo` objects track molecular types via constants `MOL_TYPE_PROTEIN`, `MOL_TYPE_DNA`, and `MOL_TYPE_RNA` to maintain chemical context.

## Code Examples

### Building a Mixed Protein-DNA-RNA Input

```python
from esm.utils.types import StructurePredictionInput, ProteinInput, DNAInput, RNAInput
from esm.models.esmfold2.processor import ESMFold2InputBuilder

# Protein chain A (single sequence, no MSA)

prot = ProteinInput(id="A", sequence="MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ", msa=None)

# DNA chain B (single strand)

dna = DNAInput(id="B", sequence="ATGCGTACGTAGCTAGCTAG")

# RNA chain C

rna = RNAInput(id="C", sequence="GGAUCCGAAUUGC")

# Assemble the full input

inp = StructurePredictionInput(sequences=[prot, dna, rna])

# Prepare tensors for the model

builder = ESMFold2InputBuilder()
features, chain_infos = builder.prepare_input(inp, device="cpu")
print("Feature keys:", list(features.keys()))
print("Chains discovered:", [c.chain_id for c in chain_infos])

```

### Inspecting Nucleotide Tokens

```python
from esm.models.esmfold2.prepare_input import build_chains_from_input

chains, tokens, atoms = build_chains_from_input(inp)
dna_chain = next(c for c in chains if c.mol_type == 1)  # MOL_TYPE_DNA = 1

print("DNA tokens (first 5):")
for t in tokens[:5]:
    if t.mol_type == 1:
        print(f"token {t.token_index}: residue {t.residue_name} atoms {t.atom_count}")

```

### Running End-to-End Structure Prediction

```python
from esm.pretrained import load_model_and_alphafold_params

# Load model and parameters

model, params = load_model_and_alphafold_params("esmfold2")
model.eval()
model.to("cuda")  # or "cpu"

# Predict structure

result = builder.fold(model, inp, device="cuda")
print("Predicted complex ID:", result.complex.id)
print("Mean pLDDT:", result.plddt.mean().item())

```

## Summary

- **Unified Input Interface**: `StructurePredictionInput` accepts any combination of `ProteinInput`, `DNAInput`, `RNAInput`, and `LigandInput` objects, enabling seamless ESMFold2 DNA RNA protein chains prediction.
- **Type-Aware Tokenization**: Nucleic acids use residue-level tokens with chemistry-specific heavy atom sets from `DNA_HEAVY_ATOMS` and `RNA_HEAVY_ATOMS`, plus backbone frames defined by C1′/C3′/C4′.
- **Hybrid Distograms**: Representative atoms for distance prediction vary by nucleotide type (C4 for purines, C2 for pyrimidines).
- **MSA Placeholders**: DNA and RNA chains receive gap-only MSAs while proteins leverage full sequence alignments.
- **End-to-End Prediction**: The pipeline in [`esm/models/esmfold2/processor.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/processor.py) processes mixed complexes in a single forward pass without external preprocessing.

## Frequently Asked Questions

### Can ESMFold2 predict DNA-protein complex structures?

Yes. ESMFold2 explicitly supports DNA-protein and RNA-protein complex prediction through its mixed-input pipeline. You pass `DNAInput` or `RNAInput` objects alongside `ProteinInput` objects to `StructurePredictionInput`, and the model processes them uniformly through the diffusion architecture.

### How does nucleotide tokenization differ from amino acid tokenization?

While standard amino acids and nucleotides both receive one token per residue, nucleotide tokens in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py) use atom lists from `DNA_HEAVY_ATOMS` or `RNA_HEAVY_ATOMS` constants. Additionally, nucleotide frames use sugar ring atoms (C1′, C3′, C4′) rather than the protein backbone (N, Cα, C).

### Does ESMFold2 use multiple sequence alignments for RNA chains?

No. According to the source code in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py), RNA and DNA chains receive gap-only MSAs where the first row contains the residue type and remaining rows are filled with `MSA_GAP_TOKEN_ID`. Only protein chains generate true MSAs for evolutionary feature extraction.

### What defines the reference frame for nucleic acid residues?

The `compute_frame_indices` function defines nucleic acid frames using three backbone atoms: **C1′, C3′, and C4′**. For distance prediction, `compute_representative_atoms` selects C4 atoms for purines (A, G) and C2 atoms for pyrimidines (C, T, U), with C1′ as the fallback for unknown nucleotides.