How to Use Pre-Installed Data Science Packages (NumPy, Pandas, scikit-learn) in LabNow AI
LabNow AI Docker images ship with NumPy, Pandas, and scikit-learn ready to import when built using the datascience Python profile, eliminating the need for manual pip installation inside running containers.
LabNow AI eliminates setup friction for data science workflows by baking essential libraries directly into its container images. According to the lab-foundation source code, these pre-installed data science packages are provisioned during the Docker build process via conditional profile arguments. This guide explains how the packages are installed and how to use them immediately upon container startup.
How LabNow AI Installs Data Science Packages Automatically
During image construction, the docker_core/Dockerfile checks the build argument ARG_PROFILE_PYTHON for the token datascience. When detected, the build appends docker_core/work/install_list_PY_datascience.pip to the installation queue.
The relevant logic appears in the Dockerfile:
&& echo "If installing Python packages" \
&& ( $(grep -q "datascience" <<< "${ARG_PROFILE_PYTHON}") && ( \
( which R && echo "rpy2 % Install rpy2 if R exists" >> /opt/utils/install_list_PY_datascience.pip || echo "Skip rpy2 install" ) \
&& ( which java && echo "py4j % Install py4j if Java exists" >> /opt/utils/install_list_PY_datascience.pip || echo "Skip py4j install" ) \
) || echo "Skip Python datascience packages install" ) \
...
&& for profile in $(echo $ARG_PROFILE_PYTHON | tr "," "\n") ; do ( \
[ -f "/opt/utils/install_list_PY_${profile}.pip" ] && install_pip "/opt/utils/install_list_PY_${profile}.pip" \
); done \
The install_pip helper function, defined in docker_core/work/script-setup.sh, executes pip install -r against the specified requirement file. This mechanism ensures NumPy, Pandas, and scikit-learn are baked into the image rather than installed at runtime.
Runtime Environment and Import Paths
The packages reside in the base Conda environment (${CONDA_PREFIX}). The Dockerfile links the Conda-installed python3.* binary to /usr/bin/python, making the pre-installed libraries available system-wide.
Because the installation occurs during the build phase, you can import these modules immediately upon entering the container:
import numpy as np
import pandas as pd
import sklearn
No virtual environment activation or additional pip install commands are required.
Practical Code Examples
The following snippets demonstrate typical usage patterns. Execute these inside any LabNow AI container without prerequisite setup steps.
NumPy Array Operations
import numpy as np
# Create a 3×3 matrix of random numbers
A = np.random.rand(3, 3)
print("Matrix A:")
print(A)
# Compute its transpose and determinant
A_T = A.T
det = np.linalg.det(A)
print("\nTranspose of A:")
print(A_T)
print(f"\nDeterminant of A = {det:.4f}")
Pandas Data Manipulation
import pandas as pd
# Load a CSV from a URL
url = "https://raw.githubusercontent.com/mwaskom/seaborn-data/master/iris.csv"
df = pd.read_csv(url)
# Display first rows
print(df.head())
# Compute mean sepal length per species
means = df.groupby("species")["sepal_length"].mean()
print("\nMean sepal length per species:")
print(means)
scikit-learn Model Training
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# Load the Iris dataset
X, y = load_iris(return_X_y=True)
# Split data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
# Train Random Forest
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)
# Evaluate
y_pred = clf.predict(X_test)
print("Accuracy:", accuracy_score(y_test, y_pred))
Key Source Files in lab-foundation
Understanding these files helps troubleshoot or extend the default package set:
docker_core/Dockerfile: Orchestrates the build and triggers profile-based installation (lines 58-63, 90-93).docker_core/work/install_list_PY_datascience.pip: Contains the explicit pip requirements includingpandasandscikit-learn.docker_core/work/script-setup.sh: Defines theinstall_pipfunction that executes the actualpip install -rcommand.
Summary
- LabNow AI pre-installs NumPy, Pandas, and scikit-learn when the
datasciencetoken is present inARG_PROFILE_PYTHON. - The installation occurs during Docker build via
docker_core/work/install_list_PY_datascience.pipand theinstall_piphelper. - Packages are available in the base Conda environment and accessible via
/usr/bin/pythonimmediately upon container startup. - No manual
pip installis required inside running containers.
Frequently Asked Questions
Do I need to run pip install for NumPy or Pandas in LabNow AI?
No. If your LabNow AI image was built with the datascience profile, these packages are already present in the global Python environment. Simply import them directly in your scripts.
How can I verify which data science packages are pre-installed?
Run pip list inside the container to view all installed packages. You can also inspect the file docker_core/work/install_list_PY_datascience.pip in the lab-foundation repository to see the intended package set before building.
Can I add additional packages to the pre-installed list?
Yes. Modify docker_core/work/install_list_PY_datascience.pip to include additional pip packages, or create a new profile file (e.g., install_list_PY_custom.pip) and add your custom profile name to the ARG_PROFILE_PYTHON build argument.
Which Python version does LabNow AI use for these packages?
LabNow AI uses the Python version installed in the base Conda environment (${CONDA_PREFIX}), which is linked to /usr/bin/python. The specific version depends on the base image tag used during the Docker build process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →