# How to Provide MSA Input to ESMC for Improved Embedding Quality

> Learn how to provide MSA input to ESMC for better protein embeddings. Use the msa parameter in ProteinInput to enhance context aware embeddings with evolutionary data.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: how-to-guide
- Published: 2026-05-30

---

**Supply a Multiple Sequence Alignment (MSA) via the `msa` parameter of `ProteinInput` when calling the ESM-C inference client to leverage evolutionary information for richer, context-aware embeddings.**

The ESM-C model in the Biohub/esm repository can utilize evolutionary context when you provide an MSA alongside your query sequence. By constructing an `MSA` object and attaching it to a `ProteinInput` dataclass, you enable the transformer architecture to attend over aligned homologs, capturing covariance signals that improve representation learning. This guide demonstrates how to construct, validate, and submit MSA data using the official SDK utilities.

## The MSA Architecture Components

ESM-C processes MSA input through a pipeline of utility classes and validation layers defined in the Biohub/esm source code.

### The ProteinInput Dataclass

The `ProteinInput` class acts as the primary container for inference requests. Defined in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py) (lines 342–347), this dataclass accepts a `sequence` string and an optional `msa` field of type `MSA`. When you instantiate `ProteinInput(sequence="MAAGKLV", msa=your_msa)`, the SDK serializes both the query sequence and the alignment data into the request payload sent to the inference endpoint.

### The MSA Utility Class

The `MSA` class, located in [`esm/utils/msa/msa.py`](https://github.com/Biohub/esm/blob/main/esm/utils/msa/msa.py) (starting at line 28), provides an object-oriented wrapper around alignment data. It stores sequences as a depth × length tensor and exposes construction methods like `MSA.from_sequences()` and `MSA.from_string()`. The class handles internal formatting, padding, and conversion to the `FastMSA` JSON representation required by the server.

### Validation and Truncation Logic

Before submission, the SDK validates MSA input through `_validate_and_truncate_msa` in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py) (lines 127–140). This function checks that the supplied object is an instance of the `MSA` class, enforces the `ESMFOLD2_MAX_MSA_SEQS` limit of 16,384 sequences, and emits warnings if the model was not trained with MSA support. If your alignment exceeds the depth limit, the function automatically truncates it to the maximum allowed size.

## Step-by-Step Implementation

Providing MSA input requires three stages: construction, packaging, and optional manual trimming.

### Constructing an MSA from Sequences

Build an `MSA` object from a list of aligned sequences (strings of equal length with gaps represented as `-`):

```python
from esm.utils.msa import MSA

aligned_seqs = [
    "MA---KLV",
    "MAAAGKLV",
    "MAVVGKLV",
    "MA---KLV",
]

msa = MSA.from_sequences(aligned_seqs)

```

### Creating the ProteinInput Object

Combine your query sequence and the MSA into a `ProteinInput` object, then pass it to the `ESMCForgeInferenceClient`:

```python
from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig, ProteinInput

protein_input = ProteinInput(
    sequence="MAAGKLV",
    msa=msa
)

client = ESMCForgeInferenceClient(model_name="esmc")
result = client.embed(
    protein_input,
    config=EmbeddingConfig(return_per_residue_embeddings=True)
)

```

### Handling Large MSAs

If your alignment contains more than 16,384 sequences, manually trim it to control which entries are retained:

```python
from esm.utils.msa import MSA

large_msa = MSA.from_fastx("my_large_msa.sto")
trimmed_msa = large_msa.select_random_sequences(16384)  # ESMFOLD2_MAX_MSA_SEQS

protein_input = ProteinInput(sequence="MAAGKLV", msa=trimmed_msa)

```

The `select_random_sequences` method is defined in [`esm/utils/msa/msa.py`](https://github.com/Biohub/esm/blob/main/esm/utils/msa/msa.py) (lines 254–260).

## Complete Code Examples

### Basic MSA Construction and Embedding Request

This example demonstrates the end-to-end workflow from construction to embedding retrieval:

```python
from esm.utils.msa import MSA
from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig, ProteinInput

# 1. Create MSA from aligned sequences

aligned_seqs = [
    "MA---KLV",
    "MAAAGKLV",
    "MAVVGKLV",
    "MA---KLV",
]
msa = MSA.from_sequences(aligned_seqs)

# 2. Build ProteinInput with sequence and MSA

protein_input = ProteinInput(
    sequence="MAAGKLV",
    msa=msa
)

# 3. Configure embedding request

embed_cfg = EmbeddingConfig(
    return_per_residue_embeddings=True,
    return_mean_embedding=True,
)

# 4. Initialize client and embed

client = ESMCForgeInferenceClient(model_name="esmc")
result = client.embed(protein_input, config=embed_cfg)

# 5. Access results

per_residue = result.per_residue_embedding  # shape: (L, D)

mean_emb = result.mean_embedding             # shape: (D,)

```

### Manual Truncation Before Submission

Control memory usage and sequence selection by trimming before SDK validation:

```python
from esm.utils.msa import MSA
from esm.sdk import ProteinInput

msa = MSA.from_fastx("large_alignment.a3m")
trimmed = msa.select_random_sequences(16384)

protein_input = ProteinInput(sequence="MAAGKLV", msa=trimmed)

```

### Using the Convenience Wrapper

The repository includes a complete runnable example in [`cookbook/snippets/esmc.py`](https://github.com/Biohub/esm/blob/main/cookbook/snippets/esmc.py) that demonstrates file import and MSA handling:

```python
from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig
from esm.widgets.utils.protein_import import import_fasta

seq, msa = import_fasta("example.fasta")
client = ESMCForgeInferenceClient()
emb = client.embed(
    protein_input=ProteinInput(sequence=seq, msa=msa),
    config=EmbeddingConfig(return_mean_embedding=True)
)

```

## How the SDK Processes MSA Data

Once you submit a `ProteinInput` containing an MSA, the SDK handles serialization and tensor preparation automatically. The `MSA` object converts to a compact `FastMSA` JSON payload for transmission. On the server side, the logic in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py) (referenced at line 1149) transforms the alignment data into tensors compatible with the ESM-C transformer. The model definition in [`esm/models/esmc.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmc.py) (forward pass starting at line 123) incorporates these MSA-derived features into its attention layers, allowing the network to attend over both the query sequence and its aligned homologs without requiring modifications to the embedding head.

## Summary

- **Construct** `MSA` objects using `MSA.from_sequences()`, `MSA.from_string()`, or `MSA.from_fastx()` utility methods defined in [`esm/utils/msa/msa.py`](https://github.com/Biohub/esm/blob/main/esm/utils/msa/msa.py).
- **Attach** the MSA to a `ProteinInput` instance via the `msa` parameter, as implemented in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py).
- **Validate** your input automatically against the `ESMFOLD2_MAX_MSA_SEQS` limit (16,384 sequences) via `_validate_and_truncate_msa` in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py).
- **Process** MSA data through the tensor preparation pipeline in [`esm/models/esmfold2/prepare_input.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmfold2/prepare_input.py) and the core model in [`esm/models/esmc.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmc.py) for improved per-residue and pooled embeddings.

## Frequently Asked Questions

### What is the maximum number of sequences allowed in an MSA for ESM-C?

The SDK enforces a hard limit of **16,384 sequences** defined by the constant `ESMFOLD2_MAX_MSA_SEQS`. If your alignment exceeds this depth, the `_validate_and_truncate_msa` function in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py) (lines 127–140) automatically truncates it. You can also manually trim alignments using the `select_random_sequences()` method from the `MSA` class in [`esm/utils/msa/msa.py`](https://github.com/Biohub/esm/blob/main/esm/utils/msa/msa.py) (lines 254–260).

### Can I provide an MSA to any ESM-C model variant?

While the `ESMCForgeInferenceClient` accepts MSA input for any inference request, the SDK emits a **warning** if the specific model weights were not trained with MSA data. According to the validation logic in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py), the request proceeds regardless, but embedding quality improvements depend on whether the model was actually trained to utilize evolutionary information.

### How does providing an MSA improve embedding quality over single-sequence input?

The transformer architecture in [`esm/models/esmc.py`](https://github.com/Biohub/esm/blob/main/esm/models/esmc.py) (forward method starting at line 123) attends over both the query sequence and aligned homologs supplied in the MSA. This allows the model to detect **evolutionary couplings** and conserved structural motifs that are invisible to single-sequence analysis, particularly improving representations for residues with low information content or remote homologs.

### What methods are available for constructing an MSA object in the ESM SDK?

According to [`esm/utils/msa/msa.py`](https://github.com/Biohub/esm/blob/main/esm/utils/msa/msa.py) (line 28 and following), you can instantiate an `MSA` using:
- `MSA.from_sequences()` with a list of aligned strings
- `MSA.from_string()` to parse formatted alignment strings
- `MSA.from_fastx()` to load from Stockholm or A3M files
The class also supports manipulation via `select_random_sequences()` and padding utilities for batch processing.