How to Specify Amino Acid Modifications Using the Modification Dataclass in StructurePredictionInput

To specify amino acid modifications in the ESM framework, instantiate Modification objects with zero-indexed positions and three-letter CCD identifiers, then assign them to the modifications list of your chain input dataclass before constructing the final StructurePredictionInput.

The esm repository by Biohub provides a structured interface for protein folding and structure prediction through the StructurePredictionInput dataclass. To model post-translational modifications or chemically altered residues, the framework exposes a dedicated Modification dataclass that integrates directly with chain-specific inputs like ProteinInput, RNAInput, and DNAInput.

Understanding the Modification Dataclass Structure

The Modification dataclass is defined in esm/utils/structure/input_builder.py (lines 16–21) and requires three fields:

  • position: An int representing the zero-indexed residue index in the sequence.
  • ccd: A str containing the three-letter Chemical Component Dictionary (CCD) identifier (e.g., "HYP" for hydroxyproline, "PSU" for pseudouridine).
  • smiles: An optional str | None field that is currently reserved and not utilized by the inference engine.

When you instantiate this class, you create a portable description of a chemical alteration that can be attached to any biological polymer chain.

Attaching Modifications to Chain Inputs

Each chain-specific dataclass—ProteinInput, RNAInput, or DNAInput—accepts a modifications parameter that takes a list of Modification objects. According to the source code in input_builder.py (lines 89–94), the serialize_structure_prediction_input function extracts these modifications and encodes them as JSON objects containing only the position and ccd fields:

if hasattr(seq_input, "modifications") and seq_input.modifications:
    mods = [
        {"position": mod.position, "ccd": mod.ccd}
        for mod in seq_input.modifications
    ]
    chain_data["modifications"] = mods

During deserialization, the private helper _mods (lines 68–73 in the same file) reconstructs the Python objects from the raw dictionary, ensuring round-trip fidelity:

def _mods(chain: dict[str, Any]) -> list[Modification] | None:
    raw = chain.get("modifications")
    if not raw:
        return None
    return [Modification(position=m["position"], ccd=m["ccd"]) for m in raw]

Practical Code Examples

Single Amino Acid Modification in a Protein

To mark residue 4 (index 3) as hydroxyproline in a protein sequence:

from esm.utils.structure.input_builder import (
    Modification,
    ProteinInput,
    StructurePredictionInput,
)

mod = Modification(position=3, ccd="HYP")
protein = ProteinInput(
    id="prot1",
    sequence="YGPKGPKGPKGKPGPDGDPGDPGDPGPKGPRG",
    modifications=[mod]
)

spi = StructurePredictionInput(sequences=[protein])

Mixed Chain Input with RNA Modifications

You can specify modifications for nucleic acids using the same dataclass. The following example adds a pseudouridine (PSU) modification at index 2 of an RNA chain:

from esm.utils.structure.input_builder import (
    Modification, ProteinInput, RNAInput, LigandInput,
    StructurePredictionInput,
)

spi = StructurePredictionInput(
    sequences=[
        ProteinInput(id="P1", sequence="MAKL"),
        RNAInput(
            id="R1",
            sequence="ACGU",
            modifications=[Modification(position=2, ccd="PSU")]
        ),
        LigandInput(id="L1", smiles="CCO")
    ]
)

Round-Trip Serialization

The SDK handles serialization automatically when sending requests to the model backend. You can also manually test the serialization pipeline:

from esm.utils.structure.input_builder import (
    serialize_structure_prediction_input,
    deserialize_structure_prediction_input,
)

payload = serialize_structure_prediction_input(spi)
restored_spi = deserialize_structure_prediction_input(payload)

assert restored_spi.sequences[1].modifications[0].ccd == "PSU"

Summary

  • The Modification dataclass in esm/utils/structure/input_builder.py defines amino acid alterations using position (zero-indexed) and ccd (three-letter code) fields.
  • Pass a list of Modification objects to the modifications parameter of ProteinInput, RNAInput, or DNAInput.
  • The serialize_structure_prediction_input function converts these objects to JSON dictionaries, while _mods handles deserialization.
  • The optional smiles field is defined but not currently consumed by the ESM inference engine.

Frequently Asked Questions

What CCD identifiers are supported for amino acid modifications?

The framework accepts any valid three-letter Chemical Component Dictionary (CCD) code recognized by the model backend, such as "HYP" for hydroxyproline or "MLY" for methyl-lysine. The Modification dataclass itself does not validate these codes; validation occurs downstream in the inference pipeline.

Can I specify modifications for RNA or DNA chains?

Yes. The modifications field exists on RNAInput and DNAInput dataclasses in addition to ProteinInput. You can use standard nucleic acid modification codes like "PSU" (pseudouridine) for RNA or "5CM" (5-methylcytosine) for DNA.

Is the optional smiles field used by the ESM inference engine?

No. While the Modification dataclass includes an optional smiles attribute for future extensibility, the current implementation in esm/utils/structure/input_builder.py only serializes the position and ccd fields, and the inference backend does not process SMILES strings for modified residues.

How does the SDK handle modification validation?

The input_builder.py module performs structural validation (ensuring the modifications list contains valid Modification objects), but it does not verify that the CCD code exists in the chemical database or that the position is within sequence bounds. The model backend raises errors if it encounters unsupported modification codes or out-of-range indices during inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →