VoiceStudio CLI vs REST API for Batch Inference: What Are the Key Differences?

The VoiceStudio CLI (omnivoice-infer) and REST API both run OmniVoice TTS inference, but differ in execution environment, batch handling, parallelism model, and result delivery.

The VoiceStudio project (debpalash/VoiceStudio) provides two interfaces for running OmniVoice text-to-speech: a command-line interface for local execution and a FastAPI-based REST API for remote access. While both leverage the same core generation logic in OmniVoice.generate, they serve different operational needs and scale differently. Understanding these differences helps you choose the right approach for your inference workload.

How the CLI and REST API Are Invoked

CLI Entry Points

The VoiceStudio CLI provides two distinct commands:

  1. omnivoice-infer — single-sample inference, implemented in omnivoice/cli/infer.py
  2. omnivoice-infer-batch — batch processing, implemented in omnivoice/cli/infer_batch.py

Both execute locally on the machine where invoked. The binary loads the model directly from a checkpoint path or Hugging Face repository and runs inference in-process.

REST API Entry Point

The REST API runs inside a FastAPI server started with uvicorn. The server loads the model once at startup and serves requests over HTTP. The primary endpoint is /v1/audio/speech, defined in omnivoice/api/v1.py.


# Start the server

uvicorn omnivoice.api.server:app --host 0.0.0.0 --port 3900

The server entry point resides in omnivoice/api/server.py, which bootstraps the FastAPI application.

Batch Handling Comparison

CLI Batch Processing

The batch CLI reads a JSON-L file (--test_list) containing multiple samples. It clusters inputs by duration or fixed batch size, then dispatches them to a pool of worker processes using ProcessPoolExecutor.


# Prepare test list

cat > test.jsonl <<EOF
{"id":"sample1","text":"Hello world!","ref_audio":"ref1.wav","ref_text":"Hello world!"}
{"id":"sample2","text":"How are you?","ref_audio":"ref2.wav","ref_text":"How are you?"}
EOF

# Run batch inference

omnivoice-infer-batch \
  --model k2-fsa/OmniVoice \
  --test_list test.jsonl \
  --res_dir results/ \
  --batch_duration 1000.0 \
  --nj_per_gpu 2

The omnivoice/cli/infer_batch.py script handles argument parsing, worker pool creation, sample clustering, and WAV file output to results/.

REST API Batch Processing

The REST API accepts a JSON payload with an array of input objects. The server groups requests internally and returns audio data for each entry.

curl -X POST http://localhost:3900/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
        "model":"k2-fsa/OmniVoice",
        "batch":[
          {"id":"sample1","text":"Hello world!","ref_audio":"/path/ref1.wav","ref_text":"Hello world!"},
          {"id":"sample2","text":"How are you?","ref_audio":"/path/ref2.wav","ref_text":"How are you?"}
        ],
        "guidance_scale":2.0,
        "num_step":32
      }' \
  -o response.json

The handler in omnivoice/api/v1.py extracts the batch list, calls OmniVoice.generate, and returns either base64-encoded WAV data or URLs to stored files.

Parallelism Model

CLI Parallelism

The CLI implements explicit multi-process parallelism through these flags:

  • --nj_per_gpu — worker processes per GPU
  • --batch_duration — target total duration per batch cluster
  • --batch_size — fixed sample count per batch

This design exploits multiple GPUs on the same node directly, with each worker maintaining its own model instance.

REST API Parallelism

Parallelism in the REST API is managed by:

  • Uvicorn worker processes (server-level)
  • Device mapping within the model
  • Request queue handling by FastAPI

Scaling typically involves launching multiple server workers rather than controlling GPU allocation per request.

Result Delivery

Aspect CLI (omnivoice-infer-batch) REST API
Output format Individual .wav files Base64-encoded audio or file URLs
Destination Local directory (--res_dir) HTTP response body
Additional output Console summary with RTF (real-time factor) JSON response with per-sample results
Streaming Not supported Supported if client requests it

After CLI execution, you'll find files like results/sample1.wav and results/sample2.wav, plus timing statistics printed to stdout. The API returns everything in a single HTTP response.

Configuration Parameters

Both interfaces accept the same generation parameters with different syntax:

CLI flags:

omnivoice-infer-batch \
  --guidance_scale 2.0 \
  --num_step 32 \
  --preprocess_prompt true

REST API JSON fields:

{
  "guidance_scale": 2.0,
  "num_step": 32,
  "preprocess_prompt": true
}

The parameter names and valid ranges are identical; only the transmission mechanism differs.

Error Handling

CLI Error Behavior

Exceptions are caught per-sample in omnivoice/cli/infer_batch.py. Failed samples are logged and the batch continues. Check console output or redirected logs for failure details.

REST API Error Behavior

HTTP status codes indicate outcomes:

  • 4xx — client errors (invalid parameters, malformed JSON)
  • 5xx — server errors (inference failures, model loading issues)

The response body contains structured error messages for programmatic handling.

When to Use Each Interface

Choose the CLI when:

  • Processing large offline corpora
  • Running scriptable pipelines
  • Direct GPU access is available
  • You need maximum throughput via ProcessPoolExecutor
  • Output files must be organized in local directories

Choose the REST API when:

  • Integrating with web or mobile applications
  • Running inference as a cloud service
  • Clients need programmatic access without local model installation
  • OpenAI-compatible client libraries are preferred
  • Request distribution across network infrastructure is required

Core Implementation Files

Component Source File Purpose
Single-sample CLI omnivoice/cli/infer.py One-shot inference with argument parsing
Batch CLI omnivoice/cli/infer_batch.py JSON-L processing, worker pools, file output
API router omnivoice/api/v1.py FastAPI routes for /v1/audio/speech
Server bootstrap omnivoice/api/server.py Uvicorn application entry point
Model definition omnivoice/models/omnivoice.py OmniVoice.from_pretrained and generate method
Audio utilities omnivoice/utils/audio.py load_audio, save_audio shared by both interfaces

These files demonstrate that OmniVoice.generate in omnivoice/models/omnivoice.py serves as the unified core, while interface-specific code handles I/O and orchestration.

Summary

  • Execution: CLI runs locally in-process; REST API serves over HTTP via FastAPI
  • Batch input: CLI uses JSON-L files; REST API accepts JSON payloads
  • Parallelism: CLI uses explicit ProcessPoolExecutor with GPU workers; REST API relies on server workers
  • Results: CLI writes .wav files to disk; REST API returns audio in HTTP responses
  • Shared core: Both use OmniVoice.generate from omnivoice/models/omnivoice.py
  • Source locations: CLI in omnivoice/cli/infer*.py, API in omnivoice/api/v1.py

Frequently Asked Questions

Can I use the REST API for the same large-scale batch jobs as the CLI?

Not optimally. The REST API handles batch requests but lacks the CLI's explicit multi-GPU process pooling. For processing thousands of samples offline, the CLI's omnivoice-infer-batch with --nj_per_gpu and --batch_duration controls provides better throughput and resource utilization. Use the API when network accessibility matters more than raw throughput.

Do the CLI and REST API produce identical audio output?

Yes, when given identical parameters. Both call the same OmniVoice.generate method in omnivoice/models/omnivoice.py. Any differences stem from parameter values, not the inference path. Verify by running the same sample through both interfaces with matching guidance_scale, num_step, and preprocess_prompt settings.

How do I choose between single-sample and batch CLI commands?

Use omnivoice-infer (from omnivoice/cli/infer.py) for one-off tests or interactive workflows. Use omnivoice-infer-batch (from omnivoice/cli/infer_batch.py) when processing multiple samples, as it implements efficient clustering and parallel execution across GPUs that single-sample invocation cannot match.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →