How scVI-tools Reference Mapping (scArches) Works for Single-Cell Data Integration

scVI-tools reference mapping uses the scArches extension to project new query datasets into a pretrained reference latent space by freezing the decoder and reinitializing only batch-specific encoder parameters, enabling rapid integration without retraining the full model.

The scArches (single-cell Architecture Surgery) framework extends scVI-tools to enable reference-based integration of new single-cell experiments. According to the K-Dense-AI/scientific-agent-skills repository, this approach stores a trained variational auto-encoder (VAE) on a large atlas and projects incoming query cells into the same latent coordinates, preserving biological relationships while correcting for technical batch effects.

The Three-Phase scArches Workflow

The implementation in scientific-skills/scvi-tools/references/workflows.md divides reference mapping into three distinct phases: training the reference model, preparing query data, and executing the mapping.

Phase 1: Training the Reference Model

First, you train a scVI variational auto-encoder on a well-annotated reference atlas. This model learns a probabilistic latent representation that captures biological variance while accounting for known batch effects.

import scvi
import scanpy as sc

# Load reference AnnData with raw counts

ref_adata = sc.read_h5ad("ref_atlas.h5ad")
scvi.model.SCVI.setup_anndata(ref_adata, batch_key="batch")

# Train the reference VAE

ref_model = scvi.model.SCVI(ref_adata)
ref_model.train()
ref_model.save("reference_scvi", overwrite=True)

The resulting model stores encoder weights, decoder weights, variational posterior parameters, and batch-correction metadata. As documented in scientific-skills/scvi-tools/references/models-specialized.md at lines 389-396, this checkpoint serves as the foundation for subsequent scArches operations.

Phase 2: Query Data Preparation

Before mapping, the query dataset must match the reference's preprocessing standards. The query AnnData must contain raw counts (not log-normalized) and use the same batch key structure as the reference.


# Load query data with identical gene panel

query_adata = sc.read_h5ad("new_experiment.h5ad")
scvi.model.SCVI.setup_anndata(query_adata, batch_key="batch")

This alignment ensures the query and reference share a common feature space, which is required for the frozen decoder to reconstruct expression correctly.

Phase 3: Reference-to-Query Mapping

The critical step uses load_query_data() to import the reference weights while preparing the model for new data. This method freezes the decoder and reinitializes only batch-specific encoder parameters, as explained in scientific-skills/scvi-tools/references/theoretical-foundations.md at lines 230-236.


# Load reference architecture and adapt for query

query_model = scvi.model.SCVI.load_query_data(
    "reference_scvi",
    adata=query_adata,
    batch_key="batch",
)

# Optional: Fine-tune only the query-specific parameters

query_model.train(max_epochs=0)

# Extract integrated latent representation

latent = query_model.get_latent_representation()

Setting max_epochs=0 skips training and projects the query immediately, while positive values enable fine-tuning of the new batch encoder adapters without touching the shared decoder weights.

Key Architectural Components

The scArches implementation relies on four technical mechanisms to ensure stable integration:

Variational Auto-Encoder (VAE): The core model learns a low-dimensional probabilistic distribution of gene expression, providing uncertainty quantification and robust batch correction through the inference network.

Frozen Decoder: By freezing decoder weights during query mapping, the model guarantees that the reference latent geometry remains unchanged. This constraint forces query cells to align to the existing biological manifold rather than distorting it.

Batch-Specific Encoder Adapters: New query batches receive their own encoder parameters (the "adapters") while sharing the decoder. This design handles novel technical effects without contaminating the reference representation.

load_query_data API: This single method handles weight loading, architecture surgery, and batch parameter initialization, eliminating error-prone manual weight copying.

Complete Code Workflow

Below is a complete pipeline combining reference training, query projection, and downstream label transfer:

import scanpy as sc
import scvi
import matplotlib.pyplot as plt

# 1. Train reference

ref_adata = sc.read_h5ad("ref_atlas.h5ad")
scvi.model.SCVI.setup_anndata(ref_adata, batch_key="batch")
ref_model = scvi.model.SCVI(ref_adata)
ref_model.train()
ref_model.save("reference_scvi", overwrite=True)

# 2. Map query

query_adata = sc.read_h5ad("query_data.h5ad")
scvi.model.SCVI.setup_anndata(query_adata, batch_key="batch")
query_model = scvi.model.SCVI.load_query_data(
    "reference_scvi",
    adata=query_adata,
    batch_key="batch",
)
query_model.train(max_epochs=0)
query_adata.obsm["X_scVI"] = query_model.get_latent_representation()

# 3. Label transfer and visualization

ref_labels = ref_adata.obs["cell_type"]
query_model.transfer_labels(ref_labels, query_adata)

# Joint UMAP

combined = ref_adata.concatenate(query_adata, batch_key="origin")
sc.pp.neighbors(combined, use_rep="X_scVI")
sc.tl.umap(combined)
sc.pl.umap(combined, color=["cell_type", "origin"])

This example assumes both datasets contain raw count matrices and share the same gene panel, which are strict requirements for scVI-tools reference mapping.

Summary

  • scVI-tools reference mapping leverages the scArches extension to integrate new single-cell datasets without retraining full variational auto-encoders.
  • The decoder remains frozen during query projection, preserving the reference latent space geometry while batch-specific encoder adapters handle new technical variation.
  • The load_query_data() method in scvi.model.SCVI automates the architecture surgery required to map queries onto pretrained references.
  • Implementation files in K-Dense-AI/scientific-agent-skills document the workflow in scientific-skills/scvi-tools/references/workflows.md and the theoretical foundations in theoretical-foundations.md.
  • Query datasets must use raw counts and match the reference gene panel exactly for valid integration.

Frequently Asked Questions

What is the difference between standard scVI and scArches reference mapping?

Standard scVI performs end-to-end training on a single dataset or integrated collection, optimizing all parameters jointly. scArches modifies this by separating reference training from query projection: it freezes the decoder and variational posterior, only optimizing new encoder parameters for incoming batches. This division allows constant-time projection of new studies without risking distortion of the reference atlas, as implemented in the load_query_data workflow described in scientific-skills/scvi-tools/references/models-specialized.md.

Why must the decoder remain frozen during reference mapping?

The decoder defines the mapping from latent space to gene expression space. Freezing these weights ensures that any point in the latent space corresponds to the same biological state across reference and query datasets. If the decoder updated during query training, the reference and query would occupy different geometric spaces, breaking integration and invalidating cell-type annotations transferred from the reference.

How do I handle gene set mismatches between reference and query datasets?

scVI-tools requires exact feature space alignment. You must subset both datasets to the intersection of genes before calling setup_anndata(). The reference mapping will fail if the query contains genes not present in the reference or vice versa. Preprocess both matrices to include identical gene lists, typically by filtering to highly variable genes identified in the reference or using a predefined marker panel.

Can I fine-tune the reference model after initial query projection?

Yes, by passing a positive value to max_epochs in query_model.train(), you enable partial fine-tuning of the query batch encoder adapters. However, the decoder and reference encoder weights remain frozen to preserve the atlas structure. This limited fine-tuning helps when the query contains severe batch effects not captured in the original training, but it should be used cautiously to avoid overfitting query-specific noise.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →