Deploying PrivateGPT Using Docker and Docker Compose: A Complete Guide
You can deploy PrivateGPT using Docker and docker-compose by cloning the zylon-ai/private-gpt repository, selecting a backend profile (Ollama CPU/CUDA/API or LlamaCPP CPU), and running docker compose --profile ollama-cpu up -d to start the containerized service on port 8001.
Deploying PrivateGPT using Docker and docker-compose provides a fully containerized, privacy-first LLM platform that requires minimal configuration. The zylon-ai/private-gpt repository ships with ready-to-use Docker images defined in Dockerfile.ollama and Dockerfile.llamacpp-cpu, along with an orchestration file that supports multiple deployment flavors for different hardware configurations.
Deployment Flavors and Architecture
PrivateGPT offers two primary deployment flavors via docker-compose.yaml, each targeting different inference backends:
- Ollama (CPU / CUDA / API): Uses the
zylonai/private-gpt:<tag>-ollamaimage and supports the profilesollama-cpu,ollama-cuda, andollama-api. This configuration includes a Traefik reverse proxy that forwards requests from port 8080 to the Ollama container on port 11434, allowing PrivateGPT to communicate with the Ollama server. - Local LlamaCPP (CPU only): Uses the
zylonai/private-gpt:<tag>-llamacpp-cpuimage with thellamacpp-cpuprofile. This runs the LlamaCPP model directly inside the PrivateGPT container without requiring external services.
Both configurations mount local directories ./local_data and ./models for persistent storage of embeddings and model files. All default configuration values live in settings-docker.yaml and can be overridden via environment variables or a .env file.
Step-by-Step Deployment Guide
1. Clone the Repository
Start by cloning the source code and navigating to the project root:
git clone https://github.com/zylon-ai/private-gpt.git
cd private-gpt
2. Configure Environment Variables
Create a .env file from the provided template to customize settings without modifying settings-docker.yaml:
cp .env.example .env
Edit .env to set secrets such as HF_TOKEN for HuggingFace model downloads, or adjust the PGPT_OLLAMA_API_BASE if using an external Ollama server.
3. Build the Images (Optional)
While Docker Compose builds images automatically on first start, you can pre-build to cache layers:
docker compose build
This command builds both the Ollama and LlamaCPP variants based on Dockerfile.ollama and Dockerfile.llamacpp-cpu.
4. Start the Stack with Compose Profiles
Select your deployment mode using the --profile flag. The docker-compose.yaml defines profile-specific services that start only the containers required for your chosen backend.
Ollama (CPU):
docker compose --profile ollama-cpu up -d
Ollama (CUDA – requires NVIDIA GPU):
docker compose --profile ollama-cuda up -d
Ollama (API only – connects to external Ollama server):
docker compose --profile ollama-api up -d
Local LlamaCPP (CPU):
docker compose --profile llamacpp-cpu up -d
Docker Compose will also start dependent services, such as the Traefik reverse proxy when using Ollama profiles.
5. Verify the Service
Once running, the PrivateGPT API exposes an OpenAI-compatible endpoint on port 8001 (or the port specified in your .env file). Verify the deployment:
curl http://localhost:8001/v1/models
You should receive a JSON response listing the available models, confirming that the service configured in settings-docker.yaml is active.
Key Configuration Files
docker-compose.yaml
This orchestration file defines the service architecture, volume mounts (./local_data, ./models), and profile groupings. It references the Dockerfiles for image builds and configures the Traefik proxy for Ollama integration.
settings-docker.yaml
Located in the repository root, this file contains default configuration values for containerized deployments, including server ports, backend modes, and model parameters. Environment variables defined in .env override these defaults at runtime.
.env.example
This template provides the structure for sensitive configuration variables including HF_TOKEN, custom ports, and API base URLs. Copy this to .env and modify values to avoid hardcoding secrets in settings-docker.yaml.
Stopping and Cleanup
To stop all running containers:
docker compose down
To stop containers and remove the persisted volumes (deleting local data and models):
docker compose down --volumes
Summary
- Clone the zylon-ai/private-gpt repository to access
docker-compose.yamland Dockerfiles - Copy
.env.exampleto.envand configure tokens, ports, or API endpoints as needed - Choose between Ollama profiles (flexible, GPU support available) or LlamaCPP CPU profile (standalone, CPU-only)
- Execute
docker compose --profile <name> up -dusing valid profiles:ollama-cpu,ollama-cuda,ollama-api, orllamacpp-cpu - Access the REST API at
http://localhost:8001and verify operation via the/v1/modelsendpoint - Modify
settings-docker.yamlor environment variables to customize runtime behavior without rebuilding images
Frequently Asked Questions
What is the difference between the Ollama and LlamaCPP deployment profiles?
The Ollama profile connects PrivateGPT to an Ollama inference server, either containerized (via the ollama service in docker-compose.yaml) or external, offering flexibility for GPU acceleration via CUDA and centralized model management. The LlamaCPP profile embeds the inference engine directly inside the PrivateGPT container using Dockerfile.llamacpp-cpu, eliminating external dependencies but restricting execution to CPU-only inference.
How do I enable GPU support when deploying PrivateGPT with Docker?
Use the ollama-cuda profile by running docker compose --profile ollama-cuda up -d, which configures the Ollama container for NVIDIA GPU access. Ensure your host has the NVIDIA Container Toolkit installed and that your zylonai/private-gpt image tag supports CUDA drivers.
Can I use an existing Ollama server instead of the containerized one?
Yes, use the ollama-api profile with docker compose --profile ollama-api up -d. This profile starts only the PrivateGPT container and the Traefik proxy, expecting you to configure PGPT_OLLAMA_API_BASE in your .env file to point to your external Ollama instance's URL (e.g., http://host.docker.internal:11434).
Where are downloaded models and vector data stored when using docker-compose?
The docker-compose.yaml mounts ./local_data and ./models from your project directory into the containers. These directories persist embeddings, chat history, and downloaded model files on your host filesystem, surviving container restarts unless you explicitly run docker compose down --volumes to remove them.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →