How to Interpret pLDDT Confidence Scores and Predicted Error in ESM‑Fold
pLDDT scores range from 0–100 where values above 90 indicate near-experimental accuracy, while the predicted aligned error (PAE) matrix quantifies pairwise residue uncertainty in Ångströms; both metrics are derived from model logits in the Biohub/esm codebase and accessible via the ProteinChain and ESMProtein APIs.
ESM‑Fold, available in the Biohub/esm repository, generates protein structure predictions together with per‑residue confidence estimates. Understanding how to interpret pLDDT confidence scores and predicted error matrices is essential for distinguishing reliable structural domains from speculative regions before downstream experimental design or docking studies.
What Are pLDDT, PAE, and TM‑Score?
ESM‑Fold outputs three distinct confidence metrics that capture different aspects of prediction quality.
pLDDT (Predicted Local Distance Difference Test)
The pLDDT is a per‑residue confidence score ranging from 0 to 100. In the codebase, these values populate the ProteinChain.confidence attribute, which is extracted from the B‑factor field during conversion from atom37 coordinates. The extraction logic resides in esm/utils/structure/protein_chain.py (lines 91–98), where chain_to_ndarray processes the raw tensor and ProteinChain.from_atom37 (lines 27–33) attaches the confidence array to the chain object.
PAE (Predicted Aligned Error)
The predicted aligned error is an L × L matrix that stores the expected RMSD between residues i and j after optimal global superposition. The implementation in esm/utils/structure/predicted_aligned_error.py (lines 37–48) converts raw model logits into expected values using bin centres (_pae_bins). Unlike pLDDT, PAE is not stored directly on the ProteinChain; it is computed on‑the‑fly from the logits tensor.
TM‑Score
The TM‑score provides a single scalar summarising global fold confidence, calculated from the PAE distribution. The function compute_tm in esm/utils/structure/predicted_aligned_error.py (lines 52–66) implements the classic Zhang‑Skolnick formula, scaling the error by d0 to produce a value between 0 and 1.
How the Code Generates These Metrics
During model inference, the ESM‑Fold decoder emits two primary tensors:
logitswith shape (L, L, N_bins) representing raw PAE logits.plddtwith shape (L,) containing per‑residue confidence values.
These are wrapped in an ESMProtein object defined in esm/sdk/api.py. When converting to a ProteinChain, the pLDDT tensor is detached and passed to ProteinChain.from_atom37 (lines 58–62 in api.py):
ProteinChain.from_atom37(..., confidence=self.plddt.detach().cpu().numpy())
For PAE, the ESMProtein.compute_predicted_aligned_error() method calls compute_predicted_aligned_error(self.logits, self.aa_mask), where aa_mask flags valid residues to handle chain breaks.
Interpreting Confidence Thresholds
Use these quantitative guidelines to filter residues and assess model utility.
pLDDT Guidelines
- > 90: Very high confidence; regions are suitable for molecular replacement or precise drug‑design applications.
- 70 – 90: Reliable backbone topology; side‑chain conformations may be uncertain.
- < 70: Low confidence; treat these regions as speculative loops or unstructured tails.
PAE Matrix Guidelines
Values are expressed in Ångströms, typically capped at ~30 Å depending on the bin configuration.
- ≤ 2 Å: Residues are positioned consistently relative to one another; suitable for interface analysis.
- > 10 Å: High uncertainty in the relative orientation; domains may be incorrectly packed.
TM‑Score Guidelines
- > 0.5: Generally indicates a correct global fold.
- > 0.8: High confidence in overall topology.
- ≈ 1.0: Near‑perfect agreement with the (unknown) native structure.
Practical Usage Examples
The following snippets assume you have installed the esm package from the Biohub/esm repository.
Extracting Per‑Residue pLDDT
Load a sequence, run inference, and access the confidence array:
from esm.sdk import ESMFold
fold = ESMFold.from_pretrained("esmfold_v1")
seq = "MKTIIALSYIFCLVFADYKDDDDK"
pred = fold(seq) # Returns ESMProtein
chain = pred.to_protein_chain() # Converts to ProteinChain
print(chain.confidence) # NumPy array, shape (L,)
The confidence attribute originates from the conversion in esm/sdk/api.py and the tensor handling in esm/utils/structure/protein_chain.py.
Computing the PAE Matrix
Generate the pairwise error matrix for visualization or domain‑packing analysis:
import matplotlib.pyplot as plt
# Compute PAE from raw logits
pae = pred.compute_predicted_aligned_error() # Shape (L, L), units Å
plt.imshow(pae.cpu().numpy(), cmap="viridis")
plt.title("Predicted Aligned Error (Å)")
plt.colorbar(label="Expected RMSD")
plt.show()
Under the hood, this invokes compute_predicted_aligned_error from esm/utils/structure/predicted_aligned_error.py, which calculates the expectation (probs * bins).sum(dim=-1) across the discretized error bins.
Calculating Global TM‑Score
Obtain a single scalar for ranking models or filtering decoys:
tm_score = pred.compute_tm_score() # Scalar between 0 and 1
print(f"TM‑score: {tm_score:.3f}")
This wraps the compute_tm function that processes the PAE logits using the standard d0 scaling factor.
Exporting Confidence for Mol*
Write an mmCIF file that embeds pLDDT as a quality metric for interactive visualization:
chain.to_mmcif("prediction_with_confidence.cif")
The to_mmcif method (lines 18–34 in esm/utils/structure/protein_chain.py) writes a ma_qa_metric block that Mol* (Molstar) recognizes for color‑by‑confidence rendering.
Summary
- pLDDT scores (0–100) reside in
ProteinChain.confidence; values above 90 mark experimental‑quality regions. - PAE matrices are generated by
ESMProtein.compute_predicted_aligned_error(), yielding pairwise expected RMSD in Ångströms. - TM‑scores derive from PAE via
ESMProtein.compute_tm_score(); exceed 0.5 for reliable global folds. - Source implementations are located in
esm/utils/structure/protein_chain.py(confidence storage),esm/sdk/api.py(API wiring), andesm/utils/structure/predicted_aligned_error.py(PAE/TM computations).
Frequently Asked Questions
What is the difference between pLDDT and PAE?
pLDDT measures local confidence for each individual residue (0–100 scale), while PAE measures the expected relative positional error between every pair of residues in Ångströms. High pLDDT means a residue is placed correctly relative to its neighbours, whereas low PAE values indicate two specific residues are positioned consistently relative to each other across the global structure.
What pLDDT threshold should I use for experimental design?
For structure‑guided drug design or molecular replacement, restrict selections to residues with pLDDT > 90. For general topology assessment or flexible loop modelling, regions scoring 70–90 are acceptable, but side‑chain rotamers should be treated as uncertain.
How do I export confidence scores for visualization in Mol*?
Use ProteinChain.to_mmcif("output.cif") to generate an mmCIF file that includes the ma_qa_metric category. Mol* automatically reads this column and can color the structure by pLDDT, showing high‑confidence regions in blue and low‑confidence regions in red or yellow according to the standard AlphaFold color scheme.
Why is my TM‑score low but my average pLDDT high?
A high average pLDDT combined with a low TM‑score usually indicates that local secondary structure elements are predicted correctly (high per‑residue confidence), but the global domain packing is wrong. Examine the PAE matrix for large off‑diagonal blocks; high values (>10 Å) between domains explain the low TM‑score despite good local accuracy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →