# How to Specify Amino Acid Modifications Using the Modification Dataclass in StructurePredictionInput

> Learn to specify amino acid modifications accurately for protein structure prediction using the Modification dataclass within StructurePredictionInput. Optimize your ESM framework workflow.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**To specify amino acid modifications in the ESM framework, instantiate `Modification` objects with zero-indexed positions and three-letter CCD identifiers, then assign them to the `modifications` list of your chain input dataclass before constructing the final `StructurePredictionInput`.**

The `esm` repository by Biohub provides a structured interface for protein folding and structure prediction through the `StructurePredictionInput` dataclass. To model post-translational modifications or chemically altered residues, the framework exposes a dedicated `Modification` dataclass that integrates directly with chain-specific inputs like `ProteinInput`, `RNAInput`, and `DNAInput`.

## Understanding the Modification Dataclass Structure

The **`Modification`** dataclass is defined in [`esm/utils/structure/input_builder.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/input_builder.py) (lines 16–21) and requires three fields:

- **`position`**: An `int` representing the zero-indexed residue index in the sequence.
- **`ccd`**: A `str` containing the three-letter Chemical Component Dictionary (CCD) identifier (e.g., `"HYP"` for hydroxyproline, `"PSU"` for pseudouridine).
- **`smiles`**: An optional `str | None` field that is currently reserved and not utilized by the inference engine.

When you instantiate this class, you create a portable description of a chemical alteration that can be attached to any biological polymer chain.

## Attaching Modifications to Chain Inputs

Each chain-specific dataclass—**`ProteinInput`**, **`RNAInput`**, or **`DNAInput`**—accepts a `modifications` parameter that takes a list of `Modification` objects. According to the source code in [`input_builder.py`](https://github.com/Biohub/esm/blob/main/input_builder.py) (lines 89–94), the `serialize_structure_prediction_input` function extracts these modifications and encodes them as JSON objects containing only the `position` and `ccd` fields:

```python
if hasattr(seq_input, "modifications") and seq_input.modifications:
    mods = [
        {"position": mod.position, "ccd": mod.ccd}
        for mod in seq_input.modifications
    ]
    chain_data["modifications"] = mods

```

During deserialization, the private helper `_mods` (lines 68–73 in the same file) reconstructs the Python objects from the raw dictionary, ensuring round-trip fidelity:

```python
def _mods(chain: dict[str, Any]) -> list[Modification] | None:
    raw = chain.get("modifications")
    if not raw:
        return None
    return [Modification(position=m["position"], ccd=m["ccd"]) for m in raw]

```

## Practical Code Examples

### Single Amino Acid Modification in a Protein

To mark residue 4 (index 3) as hydroxyproline in a protein sequence:

```python
from esm.utils.structure.input_builder import (
    Modification,
    ProteinInput,
    StructurePredictionInput,
)

mod = Modification(position=3, ccd="HYP")
protein = ProteinInput(
    id="prot1",
    sequence="YGPKGPKGPKGKPGPDGDPGDPGDPGPKGPRG",
    modifications=[mod]
)

spi = StructurePredictionInput(sequences=[protein])

```

### Mixed Chain Input with RNA Modifications

You can specify modifications for nucleic acids using the same dataclass. The following example adds a pseudouridine (**`PSU`**) modification at index 2 of an RNA chain:

```python
from esm.utils.structure.input_builder import (
    Modification, ProteinInput, RNAInput, LigandInput,
    StructurePredictionInput,
)

spi = StructurePredictionInput(
    sequences=[
        ProteinInput(id="P1", sequence="MAKL"),
        RNAInput(
            id="R1",
            sequence="ACGU",
            modifications=[Modification(position=2, ccd="PSU")]
        ),
        LigandInput(id="L1", smiles="CCO")
    ]
)

```

### Round-Trip Serialization

The SDK handles serialization automatically when sending requests to the model backend. You can also manually test the serialization pipeline:

```python
from esm.utils.structure.input_builder import (
    serialize_structure_prediction_input,
    deserialize_structure_prediction_input,
)

payload = serialize_structure_prediction_input(spi)
restored_spi = deserialize_structure_prediction_input(payload)

assert restored_spi.sequences[1].modifications[0].ccd == "PSU"

```

## Summary

- The **`Modification`** dataclass in [`esm/utils/structure/input_builder.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/input_builder.py) defines amino acid alterations using `position` (zero-indexed) and `ccd` (three-letter code) fields.
- Pass a list of `Modification` objects to the `modifications` parameter of `ProteinInput`, `RNAInput`, or `DNAInput`.
- The `serialize_structure_prediction_input` function converts these objects to JSON dictionaries, while `_mods` handles deserialization.
- The optional `smiles` field is defined but not currently consumed by the ESM inference engine.

## Frequently Asked Questions

### What CCD identifiers are supported for amino acid modifications?

The framework accepts any valid three-letter Chemical Component Dictionary (CCD) code recognized by the model backend, such as `"HYP"` for hydroxyproline or `"MLY"` for methyl-lysine. The `Modification` dataclass itself does not validate these codes; validation occurs downstream in the inference pipeline.

### Can I specify modifications for RNA or DNA chains?

Yes. The `modifications` field exists on **`RNAInput`** and **`DNAInput`** dataclasses in addition to `ProteinInput`. You can use standard nucleic acid modification codes like `"PSU"` (pseudouridine) for RNA or `"5CM"` (5-methylcytosine) for DNA.

### Is the optional `smiles` field used by the ESM inference engine?

No. While the `Modification` dataclass includes an optional `smiles` attribute for future extensibility, the current implementation in [`esm/utils/structure/input_builder.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/input_builder.py) only serializes the `position` and `ccd` fields, and the inference backend does not process SMILES strings for modified residues.

### How does the SDK handle modification validation?

The [`input_builder.py`](https://github.com/Biohub/esm/blob/main/input_builder.py) module performs structural validation (ensuring the `modifications` list contains valid `Modification` objects), but it does not verify that the CCD code exists in the chemical database or that the position is within sequence bounds. The model backend raises errors if it encounters unsupported modification codes or out-of-range indices during inference.