# MemPalace Docker Deployment with GPU Acceleration: Multi-Stage CUDA Build Guide

> Deploy MemPalace with Docker and GPU acceleration. This guide details a multi-stage CUDA build for high-performance embedding inference using onnxruntime-gpu.

- Repository: [MemPalace/mempalace](https://github.com/MemPalace/mempalace)
- Tags: how-to-guide
- Published: 2026-06-07

---

**MemPalace ships a dedicated GPU-accelerated Docker image that uses an NVIDIA CUDA base, installs the `onnxruntime-gpu` extra, and automatically configures `MEMPALACE_EMBEDDING_DEVICE=cuda` for high-performance embedding inference.**

MemPalace is an open-source memory palace framework that provides two production-ready container images. The GPU variant enables fast, on-device embeddings for large-scale mining and query workloads. This article explains the multi-stage architecture, build arguments, and runtime commands defined in the MemPalace/mempalace source code.

## MemPalace GPU Docker Image Architecture

The GPU image follows a **multi-stage build** pattern that separates the build toolchain from the runtime environment. This keeps the final image free of compilers and unnecessary development packages.

### Builder Stage

The builder starts from `nvidia/cuda:12.6.3-cudnn-runtime-ubuntu22.04` as defined by the `CUDA_IMAGE` build argument in `Dockerfile.gpu`. It uses **uv**—a fast, lock-file-driven Python installer—to resolve dependencies and create a reproducible virtual environment. The `EXTRAS` build-arg controls which optional packages are baked into the image; the default is `EXTRAS="extract,spellcheck,gpu"`, which pulls in the `onnxruntime-gpu` wheel. Because the builder stage is discarded after the build, no compilers or build-toolchain remain in the production container.

### Runtime Stage

The runtime layer copies only the pre-compiled virtual environment and a minimal interpreter from the builder. It runs as a non-root user named `mempalace` with **uid 1000**, who owns the mounted data volume. The `HOME` directory is set to `/data`, ensuring palace files, configuration, and the Hugging Face model cache (`~/.cache/huggingface`) all persist under a single volume mount. The [`docker-entrypoint.sh`](https://github.com/MemPalace/mempalace/blob/main/docker-entrypoint.sh) script forwards commands to either the MCP server (`mcp`) or any CLI sub-command.

## Building the GPU-Accelerated MemPalace Image

To build the GPU-accelerated image locally, run the following command from the repository root:

```bash
docker build -f Dockerfile.gpu -t mempalace:gpu .

```

You can customize the optional dependencies with the `EXTRAS` build argument. For example, to omit spell-check support:

```bash
docker build -f Dockerfile.gpu -t mempalace:gpu --build-arg EXTRAS="extract,gpu" .

```

If a newer `onnxruntime-gpu` wheel targets a different CUDA major version, override the `CUDA_IMAGE` argument accordingly. The `Dockerfile.gpu` contains a comment documenting this adjustment near the `ARG CUDA_IMAGE` declaration.

## Running the MemPalace GPU Container

GPU passthrough requires the **NVIDIA Container Toolkit** installed on the host. Use the following command to start the container:

```bash
docker run -i --rm --gpus all \
  -e MEMPALACE_EMBEDDING_DEVICE=cuda \
  -v mempalace-data:/data \
  mempalace:gpu

```

The `--gpus all` flag exposes host GPUs to the container. Although `Dockerfile.gpu` automatically sets `MEMPALACE_EMBEDDING_DEVICE=cuda` in the runtime image (see lines 88–93), explicitly passing the environment variable ensures visibility in your orchestration layer. The named volume `mempalace-data` persists palace files, configuration, and the downloaded embedding model cache across container restarts.

## Common GPU-Accelerated Workloads

All commands mount the same data volume so the palace state persists across runs.

- **Start the MCP server (JSON-RPC over stdio):**

```bash
docker run -i --rm --gpus all \
  -v mempalace-data:/data \
  mempalace:gpu

```

- **Run a CLI search query:**

```bash
docker run --rm --gpus all \
  -v mempalace-data:/data \
  mempalace:gpu search "why GraphQL"

```

- **Mine a local project directory:**

```bash
docker run --rm --gpus all \
  -v mempalace-data:/data \
  -v /path/to/project:/work \
  mempalace:gpu mine /work

```

In the mining example, the additional bind mount (`/path/to/project:/work`) lets the container read the target codebase without copying it into the image.

## CPU vs. GPU Image Comparison

Understanding the differences between the two variants helps you choose the right deployment target.

- **CPU image (`Dockerfile`)**: Builds from `python:3.12-slim`. It does **not** include CUDA libraries and is significantly smaller.
- **GPU image (`Dockerfile.gpu`)**: Builds from the NVIDIA CUDA runtime base. It is roughly **1.5 GB** larger but enables fast, on-device embeddings for large-scale workloads.
- **Entrypoint parity**: Both images use the same [`docker-entrypoint.sh`](https://github.com/MemPalace/mempalace/blob/main/docker-entrypoint.sh) wrapper and default to the `mcp` command.
- **Model selection**: The [`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py) module references the `MEMPALACE_EMBEDDING_DEVICE` environment variable to select between CPU and CUDA execution paths.

## Summary

- MemPalace provides a dedicated GPU image via `Dockerfile.gpu` that uses an NVIDIA CUDA base for hardware-accelerated embeddings.
- The multi-stage build uses **uv** to create a reproducible Python environment while discarding build tooling before runtime.
- The default `EXTRAS="extract,spellcheck,gpu"` installs `onnxruntime-gpu`, and `MEMPALACE_EMBEDDING_DEVICE=cuda` is preset in the runtime layer.
- Run the container with `--gpus all` and mount a persistent volume to `/data` to preserve palace state and model caches.
- Both CPU and GPU variants share the same entry-point behavior, making it easy to switch between targets without changing commands.

## Frequently Asked Questions

### What hardware and host software are required for MemPalace GPU deployment?

You need an NVIDIA GPU and the **NVIDIA Container Toolkit** installed on the Docker host. The GPU image is built on `nvidia/cuda:12.6.3-cudnn-runtime-ubuntu22.04`, so your driver must support CUDA 12.6.

### How do I switch between CPU and GPU inference without rebuilding the image?

Set the `MEMPALACE_EMBEDDING_DEVICE` environment variable to `cpu` or `cuda` at runtime. The [`mempalace/onboarding.py`](https://github.com/MemPalace/mempalace/blob/main/mempalace/onboarding.py) module reads this value to select the appropriate execution device and model cache path.

### Can I reduce the GPU image size by removing optional extras?

Yes. Pass a custom `EXTRAS` build argument when building `Dockerfile.gpu`. For example, use `--build-arg EXTRAS="gpu"` to include only the GPU dependencies and omit `extract` and `spellcheck`.

### Where does MemPalace store downloaded models inside the container?

The runtime stage sets `HOME` to `/data`, so Hugging Face caches and model weights are stored under `/data/.cache/huggingface`. Mount a named or host volume to `/data` to persist these files across container restarts.