How OMNIVOICE_MCP_OUTPUT_MODE Controls Audio Output: Inline Base64 vs. Shared Filesystem Paths

OMNIVOICE_MCP_OUTPUT_MODE is an environment variable that determines whether the VoiceStudio MCP server returns audio as inline base64-encoded data, writes it to a shared filesystem path, or both.

This configuration flag sits at the heart of the audio delivery pipeline in debpalash/VoiceStudio. By toggling a single environment variable, you control I/O behavior, network payload size, and integration patterns—critical decisions when deploying TTS services across containerized, serverless, or edge environments.

OMNIVOICE_MCP_OUTPUT_MODE Configuration Values

The MCP server recognizes three valid modes. In backend/mcp_server.py, the allowed values are strictly validated:


# backend/mcp_server.py

_OUTPUT_MODES = ("resources", "files", "both")
mode = os.environ.get("OMNIVOICE_MCP_OUTPUT_MODE", "resources").strip().lower()
if mode not in _OUTPUT_MODES:
    log.warning(
        "OMNIVOICE_MCP_OUTPUT_MODE=%r is not one of %s; using 'resources'",
        mode, _OUTPUT_MODES,
    )
    mode = "resources"

The default fallback is resources, ensuring stateless behavior when the variable is unset or invalid.

Mode 1: resources (Default) — Inline Base64

When OMNIVOICE_MCP_OUTPUT_MODE=resources, the server embeds raw WAV bytes directly in the JSON response as a base64 data URI. No filesystem I/O occurs.


# backend/mcp_server.py (excerpt)

if mode in ("resources", "both"):
    audio_data = base64.b64encode(wav_bytes).decode("ascii")
    result["audio"] = f"data:audio/wav;base64,{audio_data}"

Use this mode when:

  • Running stateless serverless functions
  • Minimizing latency for small audio clips
  • Clients expect immediate data without follow-on requests

Mode 2: files — Shared Filesystem Paths

When OMNIVOICE_MCP_OUTPUT_MODE=files, the server writes the WAV file to OUTPUTS_DIR and returns a mount-relative path in the audio_path field.


# backend/mcp_server.py (excerpt)

if mode in ("files", "both"):
    file_path = os.path.join(OUTPUTS_DIR, filename)
    with open(file_path, "wb") as f:
        f.write(wav_bytes)
    result["audio_path"] = f"/audio/{filename}"

Use this mode when:

  • Audio files must persist for downstream processing
  • Serving large files via CDN or reverse proxy
  • Operating in environments with shared volumes (Docker, Kubernetes)

The OUTPUTS_DIR is defined in backend/core/config.py and typically resolves to a subdirectory of the user data folder.

Mode 3: both — Hybrid Delivery

When OMNIVOICE_MCP_OUTPUT_MODE=both, the server executes both code paths: writing to disk and embedding base64. This provides maximum flexibility at the cost of doubled I/O and payload size.

Use this mode when:

  • Debugging or developing client integrations
  • Supporting heterogeneous clients with different capabilities
  • Migrating between deployment architectures

Practical Configuration Examples

Set the environment variable before launching the MCP server:


# Example 1 – Inline base64 only (default, stateless)

export OMNIVOICE_MCP_OUTPUT_MODE=resources
python -m omnivoice.mcp_server

# Response: {"audio": "data:audio/wav;base64,UklGRiQAAABXQVZFZm10IBAAAAABAAEAQB8AAEAfAAABAAgAZGF0YQAAAAA=..."}

# Example 2 – Files only (persistent storage)

export OMNIVOICE_MCP_OUTPUT_MODE=files
python -m omnivoice.mcp_server

# Response: {"audio_path": "/audio/tts_7a3b9f2e.wav"}

# File location: <DATA_DIR>/outputs/tts_7a3b9f2e.wav

# Example 3 – Both modes (maximum compatibility)

export OMNIVOICE_MCP_OUTPUT_MODE=both
python -m omnivoice.mcp_server

# Response: {"audio": "data:audio/wav;base64,...", "audio_path": "/audio/tts_7a3b9f2e.wav"}

Key Source Files in VoiceStudio

File Role in OMNIVOICE_MCP_OUTPUT_MODE handling
backend/mcp_server.py Reads the environment variable, validates against _OUTPUT_MODES, and branches response formatting logic
backend/core/config.py Defines OUTPUTS_DIR, the sink directory for files and both modes
tests/test_mcp_output_mode.py Validates all three modes with automated assertions

Performance and Architectural Considerations

Choosing the right OMNIVOICE_MCP_OUTPUT_MODE impacts several operational dimensions:

  • Memory vs. disk trade-off: resources keeps audio in memory (encoded as base64, ~33% size inflation), while files spills to disk
  • Network payload: Base64-encoded audio increases JSON response size by roughly 33%; large files may exceed client buffer limits
  • Container portability: files mode requires a writable volume mount at OUTPUTS_DIR; resources runs entirely in ephemeral storage
  • Caching opportunities: Files written in files or both mode can be served via HTTP caching layers; inline data must be re-encoded per request

Summary

  • OMNIVOICE_MCP_OUTPUT_MODE controls audio delivery format in VoiceStudio's MCP server through three validated values: resources, files, and both
  • The default resources mode returns inline base64 without filesystem I/O, ideal for stateless deployments
  • files mode writes to OUTPUTS_DIR and returns mount-relative paths, enabling shared storage patterns
  • both mode provides hybrid delivery for debugging and migration scenarios
  • Configuration is read once at startup from environment variables in backend/mcp_server.py

Frequently Asked Questions

What happens if OMNIVOICE_MCP_OUTPUT_MODE is set to an invalid value?

The server logs a warning and falls back to resources mode. The validation logic in backend/mcp_server.py explicitly checks against _OUTPUT_MODES = ("resources", "files", "both") and substitutes the default when the provided value does not match.

Where are audio files stored when using files or both mode?

Files are written to OUTPUTS_DIR, configured in backend/core/config.py. This directory typically resides under the application's user data folder and is exposed at the /audio URL path in API responses.

Can I switch modes without restarting the server?

No. OMNIVOICE_MCP_OUTPUT_MODE is read once at import time in backend/mcp_server.py. Changes require a process restart to take effect. For dynamic behavior, consider running multiple server instances with different mode configurations behind a load balancer.

Why does base64 encoding increase payload size?

Base64 encoding represents binary data using 64 ASCII characters, requiring 4 bytes to encode every 3 bytes of raw data—approximately a 33% overhead. This trade-off enables safe transport within JSON without additional encoding negotiations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →