# How to Interpret pLDDT Confidence Scores and Predicted Error in ESM‑Fold

> Learn to interpret pLDDT confidence scores and predicted error in ESM Fold. Understand residue uncertainty and model accuracy from Biohub/esm for precise protein structure analysis.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: deep-dive
- Published: 2026-05-30

---

**pLDDT scores range from 0–100 where values above 90 indicate near-experimental accuracy, while the predicted aligned error (PAE) matrix quantifies pairwise residue uncertainty in Ångströms; both metrics are derived from model logits in the Biohub/esm codebase and accessible via the `ProteinChain` and `ESMProtein` APIs.**

ESM‑Fold, available in the Biohub/esm repository, generates protein structure predictions together with per‑residue confidence estimates. Understanding how to interpret pLDDT confidence scores and predicted error matrices is essential for distinguishing reliable structural domains from speculative regions before downstream experimental design or docking studies.

## What Are pLDDT, PAE, and TM‑Score?

ESM‑Fold outputs three distinct confidence metrics that capture different aspects of prediction quality.

### pLDDT (Predicted Local Distance Difference Test)

The **pLDDT** is a per‑residue confidence score ranging from 0 to 100. In the codebase, these values populate the `ProteinChain.confidence` attribute, which is extracted from the B‑factor field during conversion from atom37 coordinates. The extraction logic resides in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) (lines 91–98), where `chain_to_ndarray` processes the raw tensor and `ProteinChain.from_atom37` (lines 27–33) attaches the confidence array to the chain object.

### PAE (Predicted Aligned Error)

The **predicted aligned error** is an *L × L* matrix that stores the expected RMSD between residues *i* and *j* after optimal global superposition. The implementation in [`esm/utils/structure/predicted_aligned_error.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/predicted_aligned_error.py) (lines 37–48) converts raw model logits into expected values using bin centres (`_pae_bins`). Unlike pLDDT, PAE is not stored directly on the `ProteinChain`; it is computed on‑the‑fly from the logits tensor.

### TM‑Score

The **TM‑score** provides a single scalar summarising global fold confidence, calculated from the PAE distribution. The function `compute_tm` in [`esm/utils/structure/predicted_aligned_error.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/predicted_aligned_error.py) (lines 52–66) implements the classic Zhang‑Skolnick formula, scaling the error by `d0` to produce a value between 0 and 1.

## How the Code Generates These Metrics

During model inference, the ESM‑Fold decoder emits two primary tensors:

1. `logits` with shape *(L, L, N_bins)* representing raw PAE logits.
2. `plddt` with shape *(L,)* containing per‑residue confidence values.

These are wrapped in an `ESMProtein` object defined in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py). When converting to a `ProteinChain`, the pLDDT tensor is detached and passed to `ProteinChain.from_atom37` (lines 58–62 in [`api.py`](https://github.com/Biohub/esm/blob/main/api.py)):

```python
ProteinChain.from_atom37(..., confidence=self.plddt.detach().cpu().numpy())

```

For PAE, the `ESMProtein.compute_predicted_aligned_error()` method calls `compute_predicted_aligned_error(self.logits, self.aa_mask)`, where `aa_mask` flags valid residues to handle chain breaks.

## Interpreting Confidence Thresholds

Use these quantitative guidelines to filter residues and assess model utility.

### pLDDT Guidelines

- **> 90**: Very high confidence; regions are suitable for molecular replacement or precise drug‑design applications.
- **70 – 90**: Reliable backbone topology; side‑chain conformations may be uncertain.
- **< 70**: Low confidence; treat these regions as speculative loops or unstructured tails.

### PAE Matrix Guidelines

Values are expressed in Ångströms, typically capped at ~30 Å depending on the bin configuration.

- **≤ 2 Å**: Residues are positioned consistently relative to one another; suitable for interface analysis.
- **> 10 Å**: High uncertainty in the relative orientation; domains may be incorrectly packed.

### TM‑Score Guidelines

- **> 0.5**: Generally indicates a correct global fold.
- **> 0.8**: High confidence in overall topology.
- **≈ 1.0**: Near‑perfect agreement with the (unknown) native structure.

## Practical Usage Examples

The following snippets assume you have installed the `esm` package from the Biohub/esm repository.

### Extracting Per‑Residue pLDDT

Load a sequence, run inference, and access the confidence array:

```python
from esm.sdk import ESMFold

fold = ESMFold.from_pretrained("esmfold_v1")
seq = "MKTIIALSYIFCLVFADYKDDDDK"
pred = fold(seq)                     # Returns ESMProtein

chain = pred.to_protein_chain()      # Converts to ProteinChain

print(chain.confidence)              # NumPy array, shape (L,)

```

The `confidence` attribute originates from the conversion in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) and the tensor handling in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py).

### Computing the PAE Matrix

Generate the pairwise error matrix for visualization or domain‑packing analysis:

```python
import matplotlib.pyplot as plt

# Compute PAE from raw logits

pae = pred.compute_predicted_aligned_error()   # Shape (L, L), units Å

plt.imshow(pae.cpu().numpy(), cmap="viridis")
plt.title("Predicted Aligned Error (Å)")
plt.colorbar(label="Expected RMSD")
plt.show()

```

Under the hood, this invokes `compute_predicted_aligned_error` from [`esm/utils/structure/predicted_aligned_error.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/predicted_aligned_error.py), which calculates the expectation `(probs * bins).sum(dim=-1)` across the discretized error bins.

### Calculating Global TM‑Score

Obtain a single scalar for ranking models or filtering decoys:

```python
tm_score = pred.compute_tm_score()   # Scalar between 0 and 1

print(f"TM‑score: {tm_score:.3f}")

```

This wraps the `compute_tm` function that processes the PAE logits using the standard `d0` scaling factor.

### Exporting Confidence for Mol*

Write an mmCIF file that embeds pLDDT as a quality metric for interactive visualization:

```python
chain.to_mmcif("prediction_with_confidence.cif")

```

The `to_mmcif` method (lines 18–34 in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py)) writes a `ma_qa_metric` block that Mol* (Molstar) recognizes for color‑by‑confidence rendering.

## Summary

- **pLDDT scores** (0–100) reside in `ProteinChain.confidence`; values above 90 mark experimental‑quality regions.
- **PAE matrices** are generated by `ESMProtein.compute_predicted_aligned_error()`, yielding pairwise expected RMSD in Ångströms.
- **TM‑scores** derive from PAE via `ESMProtein.compute_tm_score()`; exceed 0.5 for reliable global folds.
- Source implementations are located in [`esm/utils/structure/protein_chain.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/protein_chain.py) (confidence storage), [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) (API wiring), and [`esm/utils/structure/predicted_aligned_error.py`](https://github.com/Biohub/esm/blob/main/esm/utils/structure/predicted_aligned_error.py) (PAE/TM computations).

## Frequently Asked Questions

### What is the difference between pLDDT and PAE?

**pLDDT** measures local confidence for each individual residue (0–100 scale), while **PAE** measures the expected relative positional error between every pair of residues in Ångströms. High pLDDT means a residue is placed correctly relative to its neighbours, whereas low PAE values indicate two specific residues are positioned consistently relative to each other across the global structure.

### What pLDDT threshold should I use for experimental design?

For structure‑guided drug design or molecular replacement, restrict selections to residues with **pLDDT > 90**. For general topology assessment or flexible loop modelling, regions scoring **70–90** are acceptable, but side‑chain rotamers should be treated as uncertain.

### How do I export confidence scores for visualization in Mol*?

Use `ProteinChain.to_mmcif("output.cif")` to generate an mmCIF file that includes the `ma_qa_metric` category. Mol* automatically reads this column and can color the structure by pLDDT, showing high‑confidence regions in blue and low‑confidence regions in red or yellow according to the standard AlphaFold color scheme.

### Why is my TM‑score low but my average pLDDT high?

A high average pLDDT combined with a low TM‑score usually indicates that local secondary structure elements are predicted correctly (high per‑residue confidence), but the **global domain packing** is wrong. Examine the PAE matrix for large off‑diagonal blocks; high values (>10 Å) between domains explain the low TM‑score despite good local accuracy.