How to Set Up OpenSuperWhisper for Docker Deployment

OpenSuperWhisper is a macOS menu-bar application that you can containerize for headless Linux deployment by building the C++ inference engine and Python wrapper in a multi-stage Docker image.

OpenSuperWhisper, developed by Starmel, ships as a Swift-based macOS menu-bar utility, but its core transcription logic lives in a portable C++ library. For OpenSuperWhisper Docker deployment, you can isolate the Whisper inference engine (libwhisper) and expose it via a lightweight Python wrapper, enabling transcription in CI pipelines, servers, or any Linux environment without GUI dependencies.

Architecture Overview

The containerized setup separates the Swift UI from the inference stack. The Docker image compiles the native C++ library from libwhisper/CMakeLists.txt, installs Python dependencies defined in agent/requirements.txt, and bundles a whisper_cpp.py wrapper that loads the compiled library via ctypes.

The build process follows a multi-stage pattern:

  • Builder stage: Compiles libwhisper.a using CMake and installs Python requirements
  • Runtime stage: Copies the static library, headers, and Python wrapper into a slim Python image
  • Model layer: Includes the ggml-tiny.en.bin model file (≈100 MB) at /models, with optional bind-mount support for swapping models without rebuilding

Step 1: Create the Dockerfile

Create a Dockerfile in your project root that clones the repository and builds the C++ engine. This configuration uses a multi-stage build to keep the final image small.


# --------------------------------------------------------------

# Dockerfile – OpenSuperWhisper (Whisper-only) image

# --------------------------------------------------------------

FROM python:3.12-slim AS builder

# Install build tools & CMake

RUN apt-get update && apt-get install -y --no-install-recommends \
        build-essential cmake git && \
    rm -rf /var/lib/apt/lists/*

# ----------------------------------------------------------------

# Build the libwhisper C++ engine

# ----------------------------------------------------------------

WORKDIR /src
RUN git clone https://github.com/Starmel/OpenSuperWhisper.git . \
    && cd libwhisper \
    && cmake -B build -DCMAKE_BUILD_TYPE=Release \
    && cmake --build build --config Release

# ----------------------------------------------------------------

# Prepare the Python wrapper and install runtime deps

# ----------------------------------------------------------------

WORKDIR /app
COPY --from=builder /src/agent/requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY --from=builder /src/libwhisper/build/libwhisper.a /usr/local/lib/
COPY --from=builder /src/libwhisper/*.h /usr/local/include/

# Copy the Python wrapper

COPY --from=builder /src/agent/whisper_cpp.py /app/

# ----------------------------------------------------------------

# Runtime stage – slimmer image

# ----------------------------------------------------------------

FROM python:3.12-slim
COPY --from=builder /usr/local/lib/libwhisper.a /usr/local/lib/
COPY --from=builder /usr/local/include/whisper.h /usr/local/include/
COPY --from=builder /app/whisper_cpp.py /app/
COPY --from=builder /src/ggml-tiny.en.bin /models/ggml-tiny.en.bin

# Install runtime Python deps

COPY --from=builder /app/requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt && \
    rm requirements.txt

ENV MODEL_PATH=/models/ggml-tiny.en.bin
ENV LD_LIBRARY_PATH=/usr/local/lib

WORKDIR /app
ENTRYPOINT ["python", "whisper_cpp.py"]

The ENTRYPOINT calls the Python wrapper directly, allowing you to pass transcription arguments as command-line flags.

Step 2: Build the Image

Run the following command from the directory containing your Dockerfile:

docker build -t opensuperwhisper:latest .

This process compiles the C++ library using the CMake configuration in libwhisper/CMakeLists.txt and installs the Python dependencies listed in agent/requirements.txt, including numpy and soundfile.

Step 3: Run Transcription Jobs

Mount your audio files to /data and execute the container. The wrapper accepts the same arguments as the macOS UI's command-line options.

docker run --rm \
  -v "$(pwd)/samples:/data" \
  -v "$(pwd)/models:/models:ro" \
  opensuperwhisper:latest \
  --model /models/ggml-tiny.en.bin \
  --audio /data/audio.wav \
  --language en \
  --output /data/output.txt

Key parameters:

  • --model: Path to the Whisper model binary (default uses /models/ggml-tiny.en.bin)
  • --audio: Input audio file path inside the container
  • --language: Target language code (e.g., en, es)
  • --output: Destination for the transcription text file

Step 4: Orchestrate with Docker Compose (Optional)

For recurring jobs or integration with other services, use Docker Compose. Create a docker-compose.yml file:

version: "3.9"
services:
  whisper:
    image: opensuperwhisper:latest
    volumes:
      - ./samples:/data
      - ./models:/models:ro
    command: >
      --model /models/ggml-tiny.en.bin
      --audio /data/audio.wav
      --language en
      --output /data/output.txt

Start the transcription job with:

docker compose up

Model Management and Environment Variables

The container expects two environment variables:

  • MODEL_PATH: Default path to the Whisper model binary (set to /models/ggml-tiny.en.bin in the Dockerfile)
  • LD_LIBRARY_PATH: Set to /usr/local/lib so the Python wrapper can locate libwhisper.a via ctypes

To use a different model (e.g., ggml-base.en.bin), download the binary to your host and mount it to /models without rebuilding the image:

docker run --rm \
  -v "$(pwd)/samples:/data" \
  -v "$(pwd)/custom-models:/models:ro" \
  opensuperwhisper:latest \
  --model /models/ggml-base.en.bin \
  --audio /data/audio.wav

Summary

  • OpenSuperWhisper Docker deployment requires building the C++ inference engine from libwhisper/CMakeLists.txt and packaging the Python wrapper from the agent/ directory.
  • Use a multi-stage Dockerfile to compile libwhisper.a in a builder stage, then copy only the necessary artifacts to a slim Python runtime image.
  • Mount audio files to /data and models to /models when running containers.
  • The whisper_cpp.py wrapper loads the native library via ctypes and accepts standard Whisper CLI arguments.
  • For detailed build flags and optimization options, refer to docs/build_whisper.md in the Starmel/OpenSuperWhisper repository.

Frequently Asked Questions

Can I run OpenSuperWhisper on Linux without Docker?

OpenSuperWhisper is designed as a macOS menu-bar application written in Swift, so the full GUI application only runs on macOS. However, the core Whisper inference engine (libwhisper) is cross-platform C++ code that compiles on Linux. The Docker containerization approach described here packages that Linux-compatible inference engine with a Python wrapper, enabling headless operation on any Linux host.

How do I use a custom Whisper model in the Docker container?

The container includes the ggml-tiny.en.bin model by default at /models. To use a custom model, download the desired GGML model file (e.g., ggml-base.en.bin) to your host machine and mount it as a read-only volume to /models. Then pass the --model flag pointing to the mounted path. This approach avoids rebuilding the Docker image when switching between model sizes.

Is GPU acceleration supported in the Docker container?

The standard Dockerfile provided compiles the CPU-only version of libwhisper using the CMake configuration in libwhisper/CMakeLists.txt. To enable GPU acceleration, you would need to modify the CMake build flags (e.g., adding CUDA or Metal support) and use a base image with the appropriate GPU drivers and runtime libraries. The current implementation targets CPU inference for maximum portability across different Docker hosts.

How does the Python wrapper communicate with the C++ library?

The whisper_cpp.py wrapper (located in the agent/ directory) uses Python's built-in ctypes module to load the compiled static library libwhisper.a at runtime. It sets LD_LIBRARY_PATH to /usr/local/lib where the library is installed, then calls the C functions directly. This thin wrapper exposes the same transcription functionality used by the macOS UI but without requiring Swift or AppKit dependencies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →