VoiceStudio CLI vs REST API for Batch Inference: What Are the Key Differences?
The VoiceStudio CLI (omnivoice-infer) and REST API both run OmniVoice TTS inference, but differ in execution environment, batch handling, parallelism model, and result delivery.
The VoiceStudio project (debpalash/VoiceStudio) provides two interfaces for running OmniVoice text-to-speech: a command-line interface for local execution and a FastAPI-based REST API for remote access. While both leverage the same core generation logic in OmniVoice.generate, they serve different operational needs and scale differently. Understanding these differences helps you choose the right approach for your inference workload.
How the CLI and REST API Are Invoked
CLI Entry Points
The VoiceStudio CLI provides two distinct commands:
omnivoice-infer— single-sample inference, implemented inomnivoice/cli/infer.pyomnivoice-infer-batch— batch processing, implemented inomnivoice/cli/infer_batch.py
Both execute locally on the machine where invoked. The binary loads the model directly from a checkpoint path or Hugging Face repository and runs inference in-process.
REST API Entry Point
The REST API runs inside a FastAPI server started with uvicorn. The server loads the model once at startup and serves requests over HTTP. The primary endpoint is /v1/audio/speech, defined in omnivoice/api/v1.py.
# Start the server
uvicorn omnivoice.api.server:app --host 0.0.0.0 --port 3900
The server entry point resides in omnivoice/api/server.py, which bootstraps the FastAPI application.
Batch Handling Comparison
CLI Batch Processing
The batch CLI reads a JSON-L file (--test_list) containing multiple samples. It clusters inputs by duration or fixed batch size, then dispatches them to a pool of worker processes using ProcessPoolExecutor.
# Prepare test list
cat > test.jsonl <<EOF
{"id":"sample1","text":"Hello world!","ref_audio":"ref1.wav","ref_text":"Hello world!"}
{"id":"sample2","text":"How are you?","ref_audio":"ref2.wav","ref_text":"How are you?"}
EOF
# Run batch inference
omnivoice-infer-batch \
--model k2-fsa/OmniVoice \
--test_list test.jsonl \
--res_dir results/ \
--batch_duration 1000.0 \
--nj_per_gpu 2
The omnivoice/cli/infer_batch.py script handles argument parsing, worker pool creation, sample clustering, and WAV file output to results/.
REST API Batch Processing
The REST API accepts a JSON payload with an array of input objects. The server groups requests internally and returns audio data for each entry.
curl -X POST http://localhost:3900/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"model":"k2-fsa/OmniVoice",
"batch":[
{"id":"sample1","text":"Hello world!","ref_audio":"/path/ref1.wav","ref_text":"Hello world!"},
{"id":"sample2","text":"How are you?","ref_audio":"/path/ref2.wav","ref_text":"How are you?"}
],
"guidance_scale":2.0,
"num_step":32
}' \
-o response.json
The handler in omnivoice/api/v1.py extracts the batch list, calls OmniVoice.generate, and returns either base64-encoded WAV data or URLs to stored files.
Parallelism Model
CLI Parallelism
The CLI implements explicit multi-process parallelism through these flags:
--nj_per_gpu— worker processes per GPU--batch_duration— target total duration per batch cluster--batch_size— fixed sample count per batch
This design exploits multiple GPUs on the same node directly, with each worker maintaining its own model instance.
REST API Parallelism
Parallelism in the REST API is managed by:
- Uvicorn worker processes (server-level)
- Device mapping within the model
- Request queue handling by FastAPI
Scaling typically involves launching multiple server workers rather than controlling GPU allocation per request.
Result Delivery
| Aspect | CLI (omnivoice-infer-batch) |
REST API |
|---|---|---|
| Output format | Individual .wav files |
Base64-encoded audio or file URLs |
| Destination | Local directory (--res_dir) |
HTTP response body |
| Additional output | Console summary with RTF (real-time factor) | JSON response with per-sample results |
| Streaming | Not supported | Supported if client requests it |
After CLI execution, you'll find files like results/sample1.wav and results/sample2.wav, plus timing statistics printed to stdout. The API returns everything in a single HTTP response.
Configuration Parameters
Both interfaces accept the same generation parameters with different syntax:
CLI flags:
omnivoice-infer-batch \
--guidance_scale 2.0 \
--num_step 32 \
--preprocess_prompt true
REST API JSON fields:
{
"guidance_scale": 2.0,
"num_step": 32,
"preprocess_prompt": true
}
The parameter names and valid ranges are identical; only the transmission mechanism differs.
Error Handling
CLI Error Behavior
Exceptions are caught per-sample in omnivoice/cli/infer_batch.py. Failed samples are logged and the batch continues. Check console output or redirected logs for failure details.
REST API Error Behavior
HTTP status codes indicate outcomes:
4xx— client errors (invalid parameters, malformed JSON)5xx— server errors (inference failures, model loading issues)
The response body contains structured error messages for programmatic handling.
When to Use Each Interface
Choose the CLI when:
- Processing large offline corpora
- Running scriptable pipelines
- Direct GPU access is available
- You need maximum throughput via
ProcessPoolExecutor - Output files must be organized in local directories
Choose the REST API when:
- Integrating with web or mobile applications
- Running inference as a cloud service
- Clients need programmatic access without local model installation
- OpenAI-compatible client libraries are preferred
- Request distribution across network infrastructure is required
Core Implementation Files
| Component | Source File | Purpose |
|---|---|---|
| Single-sample CLI | omnivoice/cli/infer.py |
One-shot inference with argument parsing |
| Batch CLI | omnivoice/cli/infer_batch.py |
JSON-L processing, worker pools, file output |
| API router | omnivoice/api/v1.py |
FastAPI routes for /v1/audio/speech |
| Server bootstrap | omnivoice/api/server.py |
Uvicorn application entry point |
| Model definition | omnivoice/models/omnivoice.py |
OmniVoice.from_pretrained and generate method |
| Audio utilities | omnivoice/utils/audio.py |
load_audio, save_audio shared by both interfaces |
These files demonstrate that OmniVoice.generate in omnivoice/models/omnivoice.py serves as the unified core, while interface-specific code handles I/O and orchestration.
Summary
- Execution: CLI runs locally in-process; REST API serves over HTTP via FastAPI
- Batch input: CLI uses JSON-L files; REST API accepts JSON payloads
- Parallelism: CLI uses explicit
ProcessPoolExecutorwith GPU workers; REST API relies on server workers - Results: CLI writes
.wavfiles to disk; REST API returns audio in HTTP responses - Shared core: Both use
OmniVoice.generatefromomnivoice/models/omnivoice.py - Source locations: CLI in
omnivoice/cli/infer*.py, API inomnivoice/api/v1.py
Frequently Asked Questions
Can I use the REST API for the same large-scale batch jobs as the CLI?
Not optimally. The REST API handles batch requests but lacks the CLI's explicit multi-GPU process pooling. For processing thousands of samples offline, the CLI's omnivoice-infer-batch with --nj_per_gpu and --batch_duration controls provides better throughput and resource utilization. Use the API when network accessibility matters more than raw throughput.
Do the CLI and REST API produce identical audio output?
Yes, when given identical parameters. Both call the same OmniVoice.generate method in omnivoice/models/omnivoice.py. Any differences stem from parameter values, not the inference path. Verify by running the same sample through both interfaces with matching guidance_scale, num_step, and preprocess_prompt settings.
How do I choose between single-sample and batch CLI commands?
Use omnivoice-infer (from omnivoice/cli/infer.py) for one-off tests or interactive workflows. Use omnivoice-infer-batch (from omnivoice/cli/infer_batch.py) when processing multiple samples, as it implements efficient clustering and parallel execution across GPUs that single-sample invocation cannot match.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →