MemPalace Docker Deployment with GPU Acceleration: Multi-Stage CUDA Build Guide
MemPalace ships a dedicated GPU-accelerated Docker image that uses an NVIDIA CUDA base, installs the onnxruntime-gpu extra, and automatically configures MEMPALACE_EMBEDDING_DEVICE=cuda for high-performance embedding inference.
MemPalace is an open-source memory palace framework that provides two production-ready container images. The GPU variant enables fast, on-device embeddings for large-scale mining and query workloads. This article explains the multi-stage architecture, build arguments, and runtime commands defined in the MemPalace/mempalace source code.
MemPalace GPU Docker Image Architecture
The GPU image follows a multi-stage build pattern that separates the build toolchain from the runtime environment. This keeps the final image free of compilers and unnecessary development packages.
Builder Stage
The builder starts from nvidia/cuda:12.6.3-cudnn-runtime-ubuntu22.04 as defined by the CUDA_IMAGE build argument in Dockerfile.gpu. It uses uv—a fast, lock-file-driven Python installer—to resolve dependencies and create a reproducible virtual environment. The EXTRAS build-arg controls which optional packages are baked into the image; the default is EXTRAS="extract,spellcheck,gpu", which pulls in the onnxruntime-gpu wheel. Because the builder stage is discarded after the build, no compilers or build-toolchain remain in the production container.
Runtime Stage
The runtime layer copies only the pre-compiled virtual environment and a minimal interpreter from the builder. It runs as a non-root user named mempalace with uid 1000, who owns the mounted data volume. The HOME directory is set to /data, ensuring palace files, configuration, and the Hugging Face model cache (~/.cache/huggingface) all persist under a single volume mount. The docker-entrypoint.sh script forwards commands to either the MCP server (mcp) or any CLI sub-command.
Building the GPU-Accelerated MemPalace Image
To build the GPU-accelerated image locally, run the following command from the repository root:
docker build -f Dockerfile.gpu -t mempalace:gpu .
You can customize the optional dependencies with the EXTRAS build argument. For example, to omit spell-check support:
docker build -f Dockerfile.gpu -t mempalace:gpu --build-arg EXTRAS="extract,gpu" .
If a newer onnxruntime-gpu wheel targets a different CUDA major version, override the CUDA_IMAGE argument accordingly. The Dockerfile.gpu contains a comment documenting this adjustment near the ARG CUDA_IMAGE declaration.
Running the MemPalace GPU Container
GPU passthrough requires the NVIDIA Container Toolkit installed on the host. Use the following command to start the container:
docker run -i --rm --gpus all \
-e MEMPALACE_EMBEDDING_DEVICE=cuda \
-v mempalace-data:/data \
mempalace:gpu
The --gpus all flag exposes host GPUs to the container. Although Dockerfile.gpu automatically sets MEMPALACE_EMBEDDING_DEVICE=cuda in the runtime image (see lines 88–93), explicitly passing the environment variable ensures visibility in your orchestration layer. The named volume mempalace-data persists palace files, configuration, and the downloaded embedding model cache across container restarts.
Common GPU-Accelerated Workloads
All commands mount the same data volume so the palace state persists across runs.
- Start the MCP server (JSON-RPC over stdio):
docker run -i --rm --gpus all \
-v mempalace-data:/data \
mempalace:gpu
- Run a CLI search query:
docker run --rm --gpus all \
-v mempalace-data:/data \
mempalace:gpu search "why GraphQL"
- Mine a local project directory:
docker run --rm --gpus all \
-v mempalace-data:/data \
-v /path/to/project:/work \
mempalace:gpu mine /work
In the mining example, the additional bind mount (/path/to/project:/work) lets the container read the target codebase without copying it into the image.
CPU vs. GPU Image Comparison
Understanding the differences between the two variants helps you choose the right deployment target.
- CPU image (
Dockerfile): Builds frompython:3.12-slim. It does not include CUDA libraries and is significantly smaller. - GPU image (
Dockerfile.gpu): Builds from the NVIDIA CUDA runtime base. It is roughly 1.5 GB larger but enables fast, on-device embeddings for large-scale workloads. - Entrypoint parity: Both images use the same
docker-entrypoint.shwrapper and default to themcpcommand. - Model selection: The
mempalace/onboarding.pymodule references theMEMPALACE_EMBEDDING_DEVICEenvironment variable to select between CPU and CUDA execution paths.
Summary
- MemPalace provides a dedicated GPU image via
Dockerfile.gputhat uses an NVIDIA CUDA base for hardware-accelerated embeddings. - The multi-stage build uses uv to create a reproducible Python environment while discarding build tooling before runtime.
- The default
EXTRAS="extract,spellcheck,gpu"installsonnxruntime-gpu, andMEMPALACE_EMBEDDING_DEVICE=cudais preset in the runtime layer. - Run the container with
--gpus alland mount a persistent volume to/datato preserve palace state and model caches. - Both CPU and GPU variants share the same entry-point behavior, making it easy to switch between targets without changing commands.
Frequently Asked Questions
What hardware and host software are required for MemPalace GPU deployment?
You need an NVIDIA GPU and the NVIDIA Container Toolkit installed on the Docker host. The GPU image is built on nvidia/cuda:12.6.3-cudnn-runtime-ubuntu22.04, so your driver must support CUDA 12.6.
How do I switch between CPU and GPU inference without rebuilding the image?
Set the MEMPALACE_EMBEDDING_DEVICE environment variable to cpu or cuda at runtime. The mempalace/onboarding.py module reads this value to select the appropriate execution device and model cache path.
Can I reduce the GPU image size by removing optional extras?
Yes. Pass a custom EXTRAS build argument when building Dockerfile.gpu. For example, use --build-arg EXTRAS="gpu" to include only the GPU dependencies and omit extract and spellcheck.
Where does MemPalace store downloaded models inside the container?
The runtime stage sets HOME to /data, so Hugging Face caches and model weights are stored under /data/.cache/huggingface. Mount a named or host volume to /data to persist these files across container restarts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →