Optimizing Docker Image Size by Removing Unnecessary Packages in LabNow AI
You can reduce LabNow AI Docker images from gigabytes to hundreds of megabytes by using layered builds (atom → base → core), installing only requested profiles via build arguments, replacing system Python with Conda, and aggressively cleaning APT caches and NVIDIA build dependencies.
LabNow AI's lab-foundation repository provides a modular Docker image stack designed for data science and AI workloads. By implementing a strict layer hierarchy and conditional installation logic, the project eliminates bloatware typically found in scientific computing images. This guide explains the specific source code techniques used to strip unnecessary packages and minimize your final container size.
Layered Architecture for Minimal Base Images
The repository implements a three-tier inheritance model where each layer adds specific functionality while inheriting the minimalism of its parent:
| Layer | Purpose | Primary Dockerfile |
|---|---|---|
| atom | Minimal OS with essential utilities | docker_atom/Dockerfile |
| base | Conda installation and system Python removal | docker_base/Dockerfile (lines 14‑48) |
| core | Optional language profiles and GPU support | docker_core/Dockerfile (lines 31‑124) |
The docker_base/Dockerfile installs Miniconda via Mamba and removes the distribution's default Python. The docker_core/Dockerfile then conditionally adds language runtimes based on build arguments, ensuring you only ship the dependencies you actually need.
5 Strategies for Removing Unnecessary Packages
1. Targeted APT Installs with --no-install-recommends
The install_apt function defined in docker_atom/work/script-setup.sh explicitly passes --no-install-recommends to apt-get, preventing the installation of suggested but non-essential packages that often consume hundreds of megabytes.
# From docker_base/Dockerfile
RUN set -eux \
&& source /opt/utils/script-setup.sh \
&& install_apt /opt/utils/install_list_base.apt \
&& install__clean
2. Conditional Profile Installation via Build Arguments
Rather than including every tool by default, the core layer uses build arguments like ARG_PROFILE_PYTHON, ARG_PROFILE_R, and ARG_PROFILE_NODEJS to conditionally install packages. In docker_core/Dockerfile, the build loops through comma-separated profile values and executes only the relevant setup functions:
ARG ARG_PROFILE_NODEJS
# ...
&& for profile in $(echo $ARG_PROFILE_NODEJS | tr "," "\n") ; do ( setup_node_${profile} ) ; done \
Building with --build-arg ARG_PROFILE_PYTHON=base excludes data science, NLP, and computer vision libraries, dramatically reducing image size.
3. System Python Replacement and Removal
The base layer eliminates the operating system's Python distribution to prevent library duplication. In docker_base/Dockerfile, the build detects existing Python installations, removes the system binaries and libraries, and symlinks the Conda interpreter:
&& export SYS_PY_EXISTS=$( [ -x "$(command -v python3)" ] && echo 'true' || echo 'false' ) \
&& if $( ${SYS_PY_EXISTS:-false} && ${SYS_PY_REPLACE:-false} ) ; then \
rm -rf $(/usr/bin/python3 -c 'import sys; print(" ".join(i for i in sys.path if "python" in i))') \
&& rm -rf /usr/bin/python3* /usr/lib/python${PYTHON_VERSION} \
&& ln -sf "${CONDA_PREFIX}"/lib/python${PYTHON_VERSION} /usr/lib/ ; \
fi \
&& ln -sf "${CONDA_PREFIX}"/bin/python3.* /usr/bin/
4. Aggressive Cleanup Steps
Both the base and core layers invoke install__clean and fix_permission at the end of RUN instructions. These functions purge APT caches, temporary build files, and documentation that are not required at runtime.
# End of docker_core/Dockerfile
&& list_installed_packages && install__clean
5. Selective NVIDIA Package Removal
For GPU-enabled builds, the core Dockerfile strips massive NVIDIA Python wheels after they have been used to compile PyTorch or PaddlePaddle dependencies. Lines 95‑98 in docker_core/Dockerfile uninstall the build-time CUDA packages while retaining only the essential runtime libraries:
&& if [ "$(echo "${IDX}" | cut -c1-2)" = "cu" ] && echo "${ARG_PROFILE_PYTHON}" | grep -qE "torch|paddle" ; then \
echo "Try to uninstall nvidia python packages to reduce storage size..." \
&& pip freeze | grep -i '^nvidia-' | cut -d'=' -f1 | xargs -r pip uninstall -y \
&& apt-get -qq update --fix-missing && apt-get -qq install -y --no-install-recommends \
--allow-change-held-packages libcusparselt0 libnccl2 libnccl-dev ; \
fi \
Practical Usage Examples
Building a minimal image with only base Python:
docker build \
--build-arg BASE_NAMESPACE=labnow \
--build-arg BASE_IMG=atom \
--target base \
-t labnow/base:minimal \
-f docker_base/Dockerfile .
Building a custom core image with only the datascience profile:
docker build \
--build-arg ARG_PROFILE_PYTHON=datascience \
--build-arg ARG_PROFILE_R=base \
-t labnow/core:datascience \
-f docker_core/Dockerfile .
Summary
- Use layered builds: Start with the
atomlayer for OS utilities, addbasefor Conda, and only usecorewhen language profiles are needed. - Disable APT recommendations: The
install_aptfunction indocker_atom/work/script-setup.shuses--no-install-recommendsto avoid extraneous dependencies. - Select profiles explicitly: Pass
ARG_PROFILE_PYTHONand related arguments to install only required toolchains, skipping unnecessary ML libraries. - Replace system Python: Remove the OS Python distribution in
docker_base/Dockerfileto prevent duplicate interpreters and libraries. - Clean aggressively: Invoke
install__cleanafter installations and strip NVIDIA Python wheels after GPU compilation in the core layer.
Frequently Asked Questions
How does the atom → base → core layering reduce image size?
Each layer inherits only the minimal state of its parent. The atom layer provides a stripped-down OS, the base layer adds only Conda while removing system Python, and the core layer conditionally adds languages. This prevents the "kitchen sink" approach where a single image contains every possible tool.
Can I remove CUDA packages if I don't need GPU support?
Yes. If you omit ARG_PROFILE_PYTHON values containing torch or paddle, or build without GPU targets, the Dockerfile skips the CUDA installation logic entirely. The conditional block at lines 95‑98 of docker_core/Dockerfile only executes when building GPU-enabled images with specific ML frameworks.
What is the install__clean function and where is it defined?
The install__clean function is defined in docker_atom/work/script-setup.sh. It removes APT lists, clears caches, and deletes temporary files created during the build process. It is invoked at the end of RUN instructions in both docker_base/Dockerfile and docker_core/Dockerfile to ensure no build artifacts remain in the final image.
How do I build the smallest possible LabNow AI image?
Build only the base target with system Python replacement enabled, or build the core target with minimal profiles. Use --build-arg ARG_PROFILE_PYTHON=base and avoid GPU flags. This approach leverages the --no-install-recommends flag and cleanup functions while excluding heavy data science libraries and CUDA dependencies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →