How to Convert Predicted Structures to mmCIF Format for PDB Deposition with the Biohub/ESM Toolkit
The Biohub/esm repository provides native ProteinChain and ProteinComplex classes that export predicted structures directly to mmCIF format via the to_mmcif() and to_mmcif_string() methods, automatically embedding pLDDT confidence scores in the ma_qa_metric categories required for PDB deposition.
The Biohub/esm repository includes a comprehensive structure prediction pipeline that generates deposition-ready files from ESMFold2 outputs. Converting predicted coordinates to the mmCIF format requires careful handling of atomic data, entity metadata, and confidence metrics. This guide demonstrates how to use the native structure utilities in esm/utils/structure/ to generate validated mmCIF files containing both atomic coordinates and per-residue confidence data.
Native Structure Classes for mmCIF Export
The ESM toolkit provides two primary classes for structure representation that support direct mmCIF serialization. The ProteinChain class handles single-chain predictions, while the ProteinComplex class manages multi-chain assemblies. Both classes store atomic coordinates, residue indices, polymer sequences, and per-residue pLDDT confidence scores.
In esm/utils/structure/protein_chain.py, the ProteinChain class implements the to_mmcif() and to_mmcif_string() methods. These methods utilize Biotite's CIF writer to generate standardized mmCIF files. The implementation automatically populates the ma_qa_metric and ma_qa_metric_local categories, which are required by the PDB to store model confidence data.
Converting Single-Chain Predictions
Single-chain structures exported from ESMFold2 require minimal preprocessing before deposition. The chain object contains the atom_array, sequence, residue_index, and confidence attributes necessary for complete mmCIF generation.
Exporting to File with to_mmcif()
The to_mmcif() method writes atomic coordinates and metadata directly to disk. This approach creates a complete CIFFile object, populates it using set_structure_pdbx, and includes confidence annotations in the ma_qa_metric tables.
from esm.utils.structure.protein_chain import ProteinChain
# Assuming 'result' contains your prediction output
chain = ProteinChain.from_open_source(result)
# Write to disk with embedded confidence scores
chain.to_mmcif("predicted_structure.cif")
Generating In-Memory Strings with to_mmcif_string()
For API uploads or web deposition interfaces, use to_mmcif_string() to obtain the mmCIF content as a Python string. This method returns the same serialized content as to_mmcif() without writing to disk.
# Generate mmCIF content for direct upload
cif_text = chain.to_mmcif_string()
# Preview or transmit the structure data
print(cif_text[:500])
Handling Multi-Chain Complexes
Multi-chain predictions require entity table generation to properly define distinct polymer sequences in the mmCIF format. The ProteinComplex class in esm/utils/structure/protein_complex.py manages this through the to_mmcif_string() method and the private _add_entity_information helper.
Building and Serializing ProteinComplex Objects
Combine multiple ProteinChain instances into a ProteinComplex to generate deposition-ready files containing all chains and proper entity metadata.
from esm.utils.structure.protein_complex import ProteinComplex
# chains is a list of ProteinChain objects from multimer prediction
complex_structure = ProteinComplex.from_chains(chains)
# Generate combined mmCIF with entity tables
cif_string = complex_structure.to_mmcif_string()
with open("multimer_deposition.cif", "w") as f:
f.write(cif_string)
The implementation concatenates atom arrays from all chains, assigns unique entity IDs via _add_entity_information, and ensures that polymer sequences appear correctly in the entity and entity_poly categories.
Complete Prediction-to-Deposition Workflow
The following example demonstrates the complete pipeline from model inference to mmCIF generation using ESMFold2.
from esm.models.esmfold2 import prepare_input
from esm.pretrained import download_model
from esm.utils.structure.protein_chain import ProteinChain
import torch
# Load pretrained model
model = download_model("esmfold2")
model.eval()
# Prepare sequence input
seq = "MKTIIALSYIFCLVFADYKDDDDK"
inputs = prepare_input([seq])
# Generate prediction
with torch.no_grad():
result = model(**inputs)
# Convert to ProteinChain and export
chain = ProteinChain.from_open_source(result)
chain.to_mmcif("deposition_ready.cif")
Validating mmCIF Round-Trip Fidelity
The MmcifWrapper class in esm/utils/structure/mmcif_parsing.py provides parsing capabilities to verify deposited files. This utility reads back mmCIF files, enabling validation of atomic coordinates and confidence scores after conversion.
from esm.utils.structure.mmcif_parsing import MmcifWrapper
# Verify the generated file
wrapper = MmcifWrapper("deposition_ready.cif")
validated_structure = wrapper.get_structure()
Summary
- The
ProteinChain.to_mmcif()andto_mmcif_string()methods inesm/utils/structure/protein_chain.pyserialize single-chain predictions with embedded pLDDT scores. ProteinComplex.to_mmcif_string()inesm/utils/structure/protein_complex.pyhandles multi-chain assemblies by generating proper entity tables via_add_entity_information.- Confidence scores automatically populate the
ma_qa_metricandma_qa_metric_localcategories, satisfying PDB deposition requirements. - Biotite's
set_structure_pdbxfunction handles the underlying coordinate serialization, ensuring standard-compliant mmCIF output.
Frequently Asked Questions
What is the difference between to_mmcif and to_mmcif_string?
The to_mmcif() method accepts a file path and writes the mmCIF content directly to disk, making it suitable for local file generation. The to_mmcif_string() method returns the mmCIF content as a Python string object, which is useful for streaming uploads, API integrations, or in-memory validation before deposition.
How are confidence scores stored in the mmCIF file?
Per-residue pLDDT values are stored in the ma_qa_metric and ma_qa_metric_local categories within the mmCIF file. These categories are part of the PDBx/mmCIF dictionary extension for model quality metrics. The ProteinChain class automatically maps its internal confidence attribute to these fields during serialization in esm/utils/structure/protein_chain.py.
Can I convert multiple chains into a single mmCIF file?
Yes. Use the ProteinComplex class to aggregate multiple ProteinChain objects, then call to_mmcif_string() to generate a single mmCIF file containing all chains. The implementation in esm/utils/structure/protein_complex.py automatically creates entity tables that distinguish between different polymer sequences, which is required for proper PDB deposition of multimeric structures.
How do I validate that my mmCIF file is deposition-ready?
Use the MmcifWrapper class from esm/utils/structure/mmcif_parsing.py to read the generated file back into Python. This verifies round-trip fidelity by confirming that atomic coordinates, entity metadata, and confidence scores parse correctly. Additionally, check that the ma_qa_metric categories appear in the file header, as these are required for structures generated by prediction software.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →