How ESMFold2 Handles DNA and RNA Sequences with Protein Chains

ESMFold2 predicts structures of mixed protein-DNA-RNA complexes by converting nucleic acid sequences into unified token representations through a type-aware input pipeline that assigns entity IDs, performs nucleotide-specific tokenization, and constructs chemistry-aware reference frames, enabling end-to-end structure prediction of heterogeneous biomolecular assemblies.

ESMFold2 extends beyond protein-only folding to handle arbitrary combinations of protein, DNA, and RNA chains within a single structure prediction run. According to the Biohub/esm source code, the model achieves this capability through a sophisticated input preparation pipeline implemented primarily in esm/models/esmfold2/prepare_input.py, which treats nucleic acids as first-class citizens alongside protein polymers.

Input Normalization and Chain Construction

The process begins in esm/models/esmfold2/processor.py where the ESMFold2InputBuilder class orchestrates the preprocessing. The clean_esmfold2_input function first expands chain-break tokens (|) into separate input objects and validates that covalent-bond specifications include explicit chain IDs. This ensures that mixed complexes are correctly partitioned before tokenization.

Next, prepare_esmfold2_input calls build_chains_from_input to iterate over StructurePredictionInput.sequences. This function detects the concrete type of each entry—ProteinInput, DNAInput, or RNAInput—and assigns an entity id, asym id, and sym id to each chain. Entity deduplication ensures identical sequences share the same entity_id, while distinct asym_id values maintain separate chain instances.

Nucleotide Tokenization Strategy

In esm/models/esmfold2/prepare_input.py, the tokenize_nucleotide function handles the molecular specifics of nucleic acids. The function converts single-letter nucleotide codes to three-letter residue names (e.g., A → ADE, C → CYT) and creates one token per standard residue containing all heavy atoms defined in DNA_HEAVY_ATOMS or RNA_HEAVY_ATOMS from esm/models/esmfold2/constants.py. Modified or non-standard nucleotides fall back to atom-level tokenization, generating one token per atom.

Frame Construction and Representative Atoms

After tokenization, compute_frame_indices establishes local reference frames for each token. For nucleic acid residues, these frames utilize the sugar ring backbone atoms C1′, C3′, and C4′ (or their RNA analogs), distinct from the N-Cα-C frames used for proteins.

For distance-based structure generation, compute_representative_atoms selects specific atoms per nucleotide for distogram conditioning:

  • Purines (A, G) → C4 (with C1′ fallback)
  • Pyrimidines (C, T, U) → C2 (with C1′ fallback)
  • Unknown nucleotides → C1′

MSA Handling for Nucleic Acids

As implemented in esm/models/esmfold2/prepare_input.py, only protein chains generate true multiple-sequence alignments (MSAs). DNA and RNA chains receive gap-only MSAs where the first row contains the target residue type and subsequent rows are filled with MSA_GAP_TOKEN_ID. This satisfies the diffusion model's expectation of fixed-size tensors while bypassing evolutionary feature extraction for nucleic acids.

Tensor Assembly for Mixed Complexes

The final stage packs all per-token and per-atom data into unified tensors including coords, ref_element, ref_atom_name_chars, token_bonds, and distogram_atom_idx. These tensors contain mixed molecular types—protein, DNA, RNA, or ligands—which the diffusion model processes uniformly regardless of origin. The ChainInfo objects track molecular types via constants MOL_TYPE_PROTEIN, MOL_TYPE_DNA, and MOL_TYPE_RNA to maintain chemical context.

Code Examples

Building a Mixed Protein-DNA-RNA Input

from esm.utils.types import StructurePredictionInput, ProteinInput, DNAInput, RNAInput
from esm.models.esmfold2.processor import ESMFold2InputBuilder

# Protein chain A (single sequence, no MSA)

prot = ProteinInput(id="A", sequence="MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ", msa=None)

# DNA chain B (single strand)

dna = DNAInput(id="B", sequence="ATGCGTACGTAGCTAGCTAG")

# RNA chain C

rna = RNAInput(id="C", sequence="GGAUCCGAAUUGC")

# Assemble the full input

inp = StructurePredictionInput(sequences=[prot, dna, rna])

# Prepare tensors for the model

builder = ESMFold2InputBuilder()
features, chain_infos = builder.prepare_input(inp, device="cpu")
print("Feature keys:", list(features.keys()))
print("Chains discovered:", [c.chain_id for c in chain_infos])

Inspecting Nucleotide Tokens

from esm.models.esmfold2.prepare_input import build_chains_from_input

chains, tokens, atoms = build_chains_from_input(inp)
dna_chain = next(c for c in chains if c.mol_type == 1)  # MOL_TYPE_DNA = 1

print("DNA tokens (first 5):")
for t in tokens[:5]:
    if t.mol_type == 1:
        print(f"token {t.token_index}: residue {t.residue_name} atoms {t.atom_count}")

Running End-to-End Structure Prediction

from esm.pretrained import load_model_and_alphafold_params

# Load model and parameters

model, params = load_model_and_alphafold_params("esmfold2")
model.eval()
model.to("cuda")  # or "cpu"

# Predict structure

result = builder.fold(model, inp, device="cuda")
print("Predicted complex ID:", result.complex.id)
print("Mean pLDDT:", result.plddt.mean().item())

Summary

  • Unified Input Interface: StructurePredictionInput accepts any combination of ProteinInput, DNAInput, RNAInput, and LigandInput objects, enabling seamless ESMFold2 DNA RNA protein chains prediction.
  • Type-Aware Tokenization: Nucleic acids use residue-level tokens with chemistry-specific heavy atom sets from DNA_HEAVY_ATOMS and RNA_HEAVY_ATOMS, plus backbone frames defined by C1′/C3′/C4′.
  • Hybrid Distograms: Representative atoms for distance prediction vary by nucleotide type (C4 for purines, C2 for pyrimidines).
  • MSA Placeholders: DNA and RNA chains receive gap-only MSAs while proteins leverage full sequence alignments.
  • End-to-End Prediction: The pipeline in esm/models/esmfold2/processor.py processes mixed complexes in a single forward pass without external preprocessing.

Frequently Asked Questions

Can ESMFold2 predict DNA-protein complex structures?

Yes. ESMFold2 explicitly supports DNA-protein and RNA-protein complex prediction through its mixed-input pipeline. You pass DNAInput or RNAInput objects alongside ProteinInput objects to StructurePredictionInput, and the model processes them uniformly through the diffusion architecture.

How does nucleotide tokenization differ from amino acid tokenization?

While standard amino acids and nucleotides both receive one token per residue, nucleotide tokens in esm/models/esmfold2/prepare_input.py use atom lists from DNA_HEAVY_ATOMS or RNA_HEAVY_ATOMS constants. Additionally, nucleotide frames use sugar ring atoms (C1′, C3′, C4′) rather than the protein backbone (N, Cα, C).

Does ESMFold2 use multiple sequence alignments for RNA chains?

No. According to the source code in esm/models/esmfold2/prepare_input.py, RNA and DNA chains receive gap-only MSAs where the first row contains the residue type and remaining rows are filled with MSA_GAP_TOKEN_ID. Only protein chains generate true MSAs for evolutionary feature extraction.

What defines the reference frame for nucleic acid residues?

The compute_frame_indices function defines nucleic acid frames using three backbone atoms: C1′, C3′, and C4′. For distance prediction, compute_representative_atoms selects C4 atoms for purines (A, G) and C2 atoms for pyrimidines (C, T, U), with C1′ as the fallback for unknown nucleotides.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →