Understanding the Multi-Stage Docker Build Process in the LabNow AI Repository
The LabNow AI repository implements a four-stage Docker pipeline—docker_atom, docker_base, docker_core, and docker_docker_kit—that chains images via BASE_NAMESPACE and BASE_IMG arguments to create modular, GPU-aware containers for data science workflows.
The labnow-ai/lab-foundation repository orchestrates container construction through a sophisticated multi-stage Docker architecture designed to separate concerns across distinct image layers. This system enables incremental builds where each stage extends the previous one, allowing teams to cache intermediate images and customize language runtimes through build-time arguments without modifying the underlying Dockerfiles.
The Four-Stage Build Pipeline
The multi-stage process breaks monolithic image construction into four logical stages, each managed by its own Dockerfile and serving a specific purpose in the dependency chain.
Stage 1: docker_atom (OS Foundation)
Located at docker_atom/Dockerfile, this stage establishes the minimal operating system layer. It starts from ubuntu:noble (or a user-supplied base image) and installs essential OS packages, configures locales, and populates /opt/utils/ with utility scripts used by downstream stages. These helper scripts in docker_atom/work/*.sh provide shared functions for environment setup that later stages inherit.
Stage 2: docker_base (Python Environment)
The docker_base/Dockerfile builds upon the atom image produced in stage one. This layer installs conda and mamba, selects a specific Python version (defaulting to 3.12 via the PYTHON_VERSION argument), and optionally replaces the system Python with the conda-managed distribution. It also installs tini as an init system and performs cleanup to minimize layer size.
Stage 3: docker_core (Language Profiles)
docker_core/Dockerfile extends the base image to add comprehensive language support. It installs Node.js, Java, LaTeX, R, Go, Rust, Julia, and Octave through a profile-based system. The stage handles GPU-aware installations for TensorFlow, PyTorch, and PaddlePaddle by inspecting the presence of nvcc and the CUDA_VERSION environment variable. Profile-specific setup scripts located in docker_core/work/*.sh (such as setup_R_{profile} or setup_Python_{profile}) execute based on comma-separated ARG_PROFILE_* variables.
Stage 4: docker_docker_kit (Container Tools)
The final stage at docker_docker_kit/Dockerfile branches from the base image (stage two, not stage three) to provide Docker-specific tooling. It installs docker-compose, image-syncer, and the PyYaml Python package, creating a specialized image for container management workflows without inheriting the full language stack from core.
How Stages Connect via Build Arguments
The pipeline achieves modularity through two critical build arguments declared in every Dockerfile: ARG BASE_NAMESPACE and ARG BASE_IMG. These variables enable dynamic image resolution in the FROM instruction using the pattern ${BASE_NAMESPACE:+$BASE_NAMESPACE/}${BASE_IMG}.
When building downstream images, developers pass these arguments to link the new layer to the previously published image:
- Argument forwarding passes the namespace and image name from the previous stage
- Layer inheritance ensures the final images contain all earlier packages without reinstalling them
- Namespace flexibility supports private registries by adjusting
BASE_NAMESPACEwithout editing Dockerfiles
This chaining mechanism allows the docker_core stage to reference labnow/base:latest while docker_docker_kit can reference the same base image independently, avoiding unnecessary bloat from language profiles when only Docker tools are required.
Configurable Language Profiles and GPU Support
The docker_core stage implements a sophisticated profile system controlled by build arguments like ARG_PROFILE_PYTHON, ARG_PROFILE_R, and ARG_PROFILE_NODEJS. These accept comma-separated values (e.g., base,datascience) that determine which helper scripts execute during the build.
GPU awareness operates through conditional logic in the Dockerfile. When nvcc is detected and CUDA_VERSION is set, the build selects appropriate index URLs for machine learning libraries. The stage also removes unused NVIDIA Python packages to maintain optimal image sizes for GPU-enabled containers.
Building Images with Custom Profiles
Building the Full Stack
Execute these commands sequentially to produce the complete image hierarchy:
# Stage 1: Build the atom image
docker build \
--build-arg BASE_IMG=ubuntu:noble \
-t labnow/atom:latest \
-f docker_atom/Dockerfile .
# Stage 2: Build the base image with conda
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=atom \
-t labnow/base:latest \
-f docker_base/Dockerfile .
# Stage 3: Build the core image with multiple profiles
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=base \
--build-arg ARG_PROFILE_PYTHON=base,datascience \
--build-arg ARG_PROFILE_R=base \
--build-arg ARG_PROFILE_NODEJS=node \
-t labnow/core:latest \
-f docker_core/Dockerfile .
# Stage 4: Build the Docker toolkit
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=base \
-t labnow/docker-kit:latest \
-f docker_docker_kit/Dockerfile .
Creating a Lightweight CPU-Only Python Image
To build a minimal image containing only base Python utilities without GPU dependencies or additional languages:
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=base \
--build-arg ARG_PROFILE_PYTHON=base \
-t labnow/python-cpu:latest \
-f docker_core/Dockerfile .
Overriding the Python Version
Specify an alternative Python version during the base stage construction:
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=atom \
--build-arg PYTHON_VERSION=3.11 \
-t labnow/python-3.11:latest \
-f docker_base/Dockerfile .
Summary
- Modularity: Each stage isolates distinct concerns—OS utilities, Python environment, language stacks, and Docker tooling—into separate images.
- Reusability: Intermediate images (
atom,base,core) serve as cached foundations for independent workflows, such as CI testing requiring only the Python stack. - Configurability: Build-time arguments enable CPU-only, GPU-enabled, or minimal installations from identical Dockerfiles without source modifications.
- Maintainability: Adding new language support requires only a new profile script in
docker_core/work/and an entry in the correspondingARG_PROFILE_*variable.
Frequently Asked Questions
How do I build only a specific stage of the LabNow AI Docker pipeline?
Run the docker build command targeting the specific Dockerfile and passing the appropriate BASE_NAMESPACE and BASE_IMG arguments to link to existing upstream images. For example, to build only the docker_core stage, ensure the labnow/base image exists locally or in your registry, then execute the build command with --build-arg BASE_IMG=base.
What is the difference between the base and core images?
The base image (produced by docker_base/Dockerfile) contains only the operating system utilities and conda-managed Python environment. The core image (produced by docker_core/Dockerfile) extends base by adding language-specific profiles for Node.js, R, Java, and other runtimes, plus GPU-aware machine learning libraries.
How does the build process handle GPU dependencies?
The docker_core Dockerfile inspects the build environment for nvcc and the CUDA_VERSION environment variable. When detected, it configures pip to use CUDA-specific index URLs for installing TensorFlow, PyTorch, and PaddlePaddle. The build process also removes unnecessary NVIDIA Python packages to reduce the final image size for GPU-enabled containers.
Can I add a new programming language to the core image?
Yes. Create a new setup script in docker_core/work/ following the naming convention setup_{LANGUAGE}_{profile}.sh, then pass the profile name via the corresponding build argument (e.g., --build-arg ARG_PROFILE_NEWLANG=customprofile). The docker_core Dockerfile iterates over these arguments and executes matching scripts automatically without requiring modifications to the Dockerfile itself.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →