# How to Convert Predicted Structures to mmCIF Format for PDB Deposition with the Biohub/ESM Toolkit

> Easily convert predicted structures to mmCIF format for PDB deposition using the Biohub/esm toolkit. Export directly with confidence scores embedded for seamless submission.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**The Biohub/esm repository provides native `ProteinChain` and `ProteinComplex` classes that export predicted structures directly to mmCIF format via the `to_mmcif()` and `to_mmcif_string()` methods, automatically embedding pLDDT confidence scores in the `ma_qa_metric` categories required for PDB deposition.**

The Biohub/esm repository includes a comprehensive structure prediction pipeline that generates deposition-ready files from ESMFold2 outputs. Converting predicted coordinates to the mmCIF format requires careful handling of atomic data, entity metadata, and confidence metrics. This guide demonstrates how to use the native structure utilities in `esm/utils/structure/` to generate validated mmCIF files containing both atomic coordinates and per-residue confidence data.

## Native Structure Classes for mmCIF Export

The ESM toolkit provides two primary classes for structure representation that support direct mmCIF serialization. The **`ProteinChain`** class handles single-chain predictions, while the **`ProteinComplex`** class manages multi-chain assemblies. Both classes store atomic coordinates, residue indices, polymer sequences, and per-residue pLDDT confidence scores.

In [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py), the `ProteinChain` class implements the **`to_mmcif()`** and **`to_mmcif_string()`** methods. These methods utilize Biotite's CIF writer to generate standardized mmCIF files. The implementation automatically populates the **`ma_qa_metric`** and **`ma_qa_metric_local`** categories, which are required by the PDB to store model confidence data.

## Converting Single-Chain Predictions

Single-chain structures exported from ESMFold2 require minimal preprocessing before deposition. The chain object contains the **`atom_array`**, **`sequence`**, **`residue_index`**, and **`confidence`** attributes necessary for complete mmCIF generation.

### Exporting to File with to_mmcif()

The **`to_mmcif()`** method writes atomic coordinates and metadata directly to disk. This approach creates a complete `CIFFile` object, populates it using `set_structure_pdbx`, and includes confidence annotations in the `ma_qa_metric` tables.

```python
from esm.utils.structure.protein_chain import ProteinChain

# Assuming 'result' contains your prediction output

chain = ProteinChain.from_open_source(result)

# Write to disk with embedded confidence scores

chain.to_mmcif("predicted_structure.cif")

```

### Generating In-Memory Strings with to_mmcif_string()

For API uploads or web deposition interfaces, use **`to_mmcif_string()`** to obtain the mmCIF content as a Python string. This method returns the same serialized content as `to_mmcif()` without writing to disk.

```python

# Generate mmCIF content for direct upload

cif_text = chain.to_mmcif_string()

# Preview or transmit the structure data

print(cif_text[:500])

```

## Handling Multi-Chain Complexes

Multi-chain predictions require entity table generation to properly define distinct polymer sequences in the mmCIF format. The `ProteinComplex` class in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) manages this through the **`to_mmcif_string()`** method and the private **`_add_entity_information`** helper.

### Building and Serializing ProteinComplex Objects

Combine multiple `ProteinChain` instances into a `ProteinComplex` to generate deposition-ready files containing all chains and proper entity metadata.

```python
from esm.utils.structure.protein_complex import ProteinComplex

# chains is a list of ProteinChain objects from multimer prediction

complex_structure = ProteinComplex.from_chains(chains)

# Generate combined mmCIF with entity tables

cif_string = complex_structure.to_mmcif_string()

with open("multimer_deposition.cif", "w") as f:
    f.write(cif_string)

```

The implementation concatenates atom arrays from all chains, assigns unique entity IDs via `_add_entity_information`, and ensures that polymer sequences appear correctly in the `entity` and `entity_poly` categories.

## Complete Prediction-to-Deposition Workflow

The following example demonstrates the complete pipeline from model inference to mmCIF generation using ESMFold2.

```python
from esm.models.esmfold2 import prepare_input
from esm.pretrained import download_model
from esm.utils.structure.protein_chain import ProteinChain
import torch

# Load pretrained model

model = download_model("esmfold2")
model.eval()

# Prepare sequence input

seq = "MKTIIALSYIFCLVFADYKDDDDK"
inputs = prepare_input([seq])

# Generate prediction

with torch.no_grad():
    result = model(**inputs)

# Convert to ProteinChain and export

chain = ProteinChain.from_open_source(result)
chain.to_mmcif("deposition_ready.cif")

```

## Validating mmCIF Round-Trip Fidelity

The **`MmcifWrapper`** class in [`esm/utils/structure/mmcif_parsing.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/mmcif_parsing.py) provides parsing capabilities to verify deposited files. This utility reads back mmCIF files, enabling validation of atomic coordinates and confidence scores after conversion.

```python
from esm.utils.structure.mmcif_parsing import MmcifWrapper

# Verify the generated file

wrapper = MmcifWrapper("deposition_ready.cif")
validated_structure = wrapper.get_structure()

```

## Summary

- The **`ProteinChain.to_mmcif()`** and **`to_mmcif_string()`** methods in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) serialize single-chain predictions with embedded pLDDT scores.
- **`ProteinComplex.to_mmcif_string()`** in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) handles multi-chain assemblies by generating proper entity tables via `_add_entity_information`.
- Confidence scores automatically populate the **`ma_qa_metric`** and **`ma_qa_metric_local`** categories, satisfying PDB deposition requirements.
- Biotite's `set_structure_pdbx` function handles the underlying coordinate serialization, ensuring standard-compliant mmCIF output.

## Frequently Asked Questions

### What is the difference between to_mmcif and to_mmcif_string?

The **`to_mmcif()`** method accepts a file path and writes the mmCIF content directly to disk, making it suitable for local file generation. The **`to_mmcif_string()`** method returns the mmCIF content as a Python string object, which is useful for streaming uploads, API integrations, or in-memory validation before deposition.

### How are confidence scores stored in the mmCIF file?

Per-residue pLDDT values are stored in the **`ma_qa_metric`** and **`ma_qa_metric_local`** categories within the mmCIF file. These categories are part of the PDBx/mmCIF dictionary extension for model quality metrics. The `ProteinChain` class automatically maps its internal `confidence` attribute to these fields during serialization in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py).

### Can I convert multiple chains into a single mmCIF file?

Yes. Use the **`ProteinComplex`** class to aggregate multiple `ProteinChain` objects, then call **`to_mmcif_string()`** to generate a single mmCIF file containing all chains. The implementation in [`esm/utils/structure/protein_complex.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_complex.py) automatically creates entity tables that distinguish between different polymer sequences, which is required for proper PDB deposition of multimeric structures.

### How do I validate that my mmCIF file is deposition-ready?

Use the **`MmcifWrapper`** class from [`esm/utils/structure/mmcif_parsing.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/mmcif_parsing.py) to read the generated file back into Python. This verifies round-trip fidelity by confirming that atomic coordinates, entity metadata, and confidence scores parse correctly. Additionally, check that the `ma_qa_metric` categories appear in the file header, as these are required for structures generated by prediction software.