How to Provide MSA Input to ESMC for Improved Embedding Quality

Supply a Multiple Sequence Alignment (MSA) via the msa parameter of ProteinInput when calling the ESM-C inference client to leverage evolutionary information for richer, context-aware embeddings.

The ESM-C model in the Biohub/esm repository can utilize evolutionary context when you provide an MSA alongside your query sequence. By constructing an MSA object and attaching it to a ProteinInput dataclass, you enable the transformer architecture to attend over aligned homologs, capturing covariance signals that improve representation learning. This guide demonstrates how to construct, validate, and submit MSA data using the official SDK utilities.

The MSA Architecture Components

ESM-C processes MSA input through a pipeline of utility classes and validation layers defined in the Biohub/esm source code.

The ProteinInput Dataclass

The ProteinInput class acts as the primary container for inference requests. Defined in esm/sdk/forge.py (lines 342–347), this dataclass accepts a sequence string and an optional msa field of type MSA. When you instantiate ProteinInput(sequence="MAAGKLV", msa=your_msa), the SDK serializes both the query sequence and the alignment data into the request payload sent to the inference endpoint.

The MSA Utility Class

The MSA class, located in esm/utils/msa/msa.py (starting at line 28), provides an object-oriented wrapper around alignment data. It stores sequences as a depth × length tensor and exposes construction methods like MSA.from_sequences() and MSA.from_string(). The class handles internal formatting, padding, and conversion to the FastMSA JSON representation required by the server.

Validation and Truncation Logic

Before submission, the SDK validates MSA input through _validate_and_truncate_msa in esm/sdk/forge.py (lines 127–140). This function checks that the supplied object is an instance of the MSA class, enforces the ESMFOLD2_MAX_MSA_SEQS limit of 16,384 sequences, and emits warnings if the model was not trained with MSA support. If your alignment exceeds the depth limit, the function automatically truncates it to the maximum allowed size.

Step-by-Step Implementation

Providing MSA input requires three stages: construction, packaging, and optional manual trimming.

Constructing an MSA from Sequences

Build an MSA object from a list of aligned sequences (strings of equal length with gaps represented as -):

from esm.utils.msa import MSA

aligned_seqs = [
    "MA---KLV",
    "MAAAGKLV",
    "MAVVGKLV",
    "MA---KLV",
]

msa = MSA.from_sequences(aligned_seqs)

Creating the ProteinInput Object

Combine your query sequence and the MSA into a ProteinInput object, then pass it to the ESMCForgeInferenceClient:

from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig, ProteinInput

protein_input = ProteinInput(
    sequence="MAAGKLV",
    msa=msa
)

client = ESMCForgeInferenceClient(model_name="esmc")
result = client.embed(
    protein_input,
    config=EmbeddingConfig(return_per_residue_embeddings=True)
)

Handling Large MSAs

If your alignment contains more than 16,384 sequences, manually trim it to control which entries are retained:

from esm.utils.msa import MSA

large_msa = MSA.from_fastx("my_large_msa.sto")
trimmed_msa = large_msa.select_random_sequences(16384)  # ESMFOLD2_MAX_MSA_SEQS

protein_input = ProteinInput(sequence="MAAGKLV", msa=trimmed_msa)

The select_random_sequences method is defined in esm/utils/msa/msa.py (lines 254–260).

Complete Code Examples

Basic MSA Construction and Embedding Request

This example demonstrates the end-to-end workflow from construction to embedding retrieval:

from esm.utils.msa import MSA
from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig, ProteinInput

# 1. Create MSA from aligned sequences

aligned_seqs = [
    "MA---KLV",
    "MAAAGKLV",
    "MAVVGKLV",
    "MA---KLV",
]
msa = MSA.from_sequences(aligned_seqs)

# 2. Build ProteinInput with sequence and MSA

protein_input = ProteinInput(
    sequence="MAAGKLV",
    msa=msa
)

# 3. Configure embedding request

embed_cfg = EmbeddingConfig(
    return_per_residue_embeddings=True,
    return_mean_embedding=True,
)

# 4. Initialize client and embed

client = ESMCForgeInferenceClient(model_name="esmc")
result = client.embed(protein_input, config=embed_cfg)

# 5. Access results

per_residue = result.per_residue_embedding  # shape: (L, D)

mean_emb = result.mean_embedding             # shape: (D,)

Manual Truncation Before Submission

Control memory usage and sequence selection by trimming before SDK validation:

from esm.utils.msa import MSA
from esm.sdk import ProteinInput

msa = MSA.from_fastx("large_alignment.a3m")
trimmed = msa.select_random_sequences(16384)

protein_input = ProteinInput(sequence="MAAGKLV", msa=trimmed)

Using the Convenience Wrapper

The repository includes a complete runnable example in cookbook/snippets/esmc.py that demonstrates file import and MSA handling:

from esm.sdk import ESMCForgeInferenceClient, EmbeddingConfig
from esm.widgets.utils.protein_import import import_fasta

seq, msa = import_fasta("example.fasta")
client = ESMCForgeInferenceClient()
emb = client.embed(
    protein_input=ProteinInput(sequence=seq, msa=msa),
    config=EmbeddingConfig(return_mean_embedding=True)
)

How the SDK Processes MSA Data

Once you submit a ProteinInput containing an MSA, the SDK handles serialization and tensor preparation automatically. The MSA object converts to a compact FastMSA JSON payload for transmission. On the server side, the logic in esm/models/esmfold2/prepare_input.py (referenced at line 1149) transforms the alignment data into tensors compatible with the ESM-C transformer. The model definition in esm/models/esmc.py (forward pass starting at line 123) incorporates these MSA-derived features into its attention layers, allowing the network to attend over both the query sequence and its aligned homologs without requiring modifications to the embedding head.

Summary

  • Construct MSA objects using MSA.from_sequences(), MSA.from_string(), or MSA.from_fastx() utility methods defined in esm/utils/msa/msa.py.
  • Attach the MSA to a ProteinInput instance via the msa parameter, as implemented in esm/sdk/forge.py.
  • Validate your input automatically against the ESMFOLD2_MAX_MSA_SEQS limit (16,384 sequences) via _validate_and_truncate_msa in esm/sdk/forge.py.
  • Process MSA data through the tensor preparation pipeline in esm/models/esmfold2/prepare_input.py and the core model in esm/models/esmc.py for improved per-residue and pooled embeddings.

Frequently Asked Questions

What is the maximum number of sequences allowed in an MSA for ESM-C?

The SDK enforces a hard limit of 16,384 sequences defined by the constant ESMFOLD2_MAX_MSA_SEQS. If your alignment exceeds this depth, the _validate_and_truncate_msa function in esm/sdk/forge.py (lines 127–140) automatically truncates it. You can also manually trim alignments using the select_random_sequences() method from the MSA class in esm/utils/msa/msa.py (lines 254–260).

Can I provide an MSA to any ESM-C model variant?

While the ESMCForgeInferenceClient accepts MSA input for any inference request, the SDK emits a warning if the specific model weights were not trained with MSA data. According to the validation logic in esm/sdk/forge.py, the request proceeds regardless, but embedding quality improvements depend on whether the model was actually trained to utilize evolutionary information.

How does providing an MSA improve embedding quality over single-sequence input?

The transformer architecture in esm/models/esmc.py (forward method starting at line 123) attends over both the query sequence and aligned homologs supplied in the MSA. This allows the model to detect evolutionary couplings and conserved structural motifs that are invisible to single-sequence analysis, particularly improving representations for residues with low information content or remote homologs.

What methods are available for constructing an MSA object in the ESM SDK?

According to esm/utils/msa/msa.py (line 28 and following), you can instantiate an MSA using:

  • MSA.from_sequences() with a list of aligned strings
  • MSA.from_string() to parse formatted alignment strings
  • MSA.from_fastx() to load from Stockholm or A3M files The class also supports manipulation via select_random_sequences() and padding utilities for batch processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →