# VoiceStudio CLI vs REST API for Batch Inference: What Are the Key Differences?

> Compare VoiceStudio CLI and REST API for batch inference. Understand key differences in environment, batching, parallelism, and result delivery for efficient TTS.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: comparison
- Published: 2026-09-06

---

**The VoiceStudio CLI (`omnivoice-infer`) and REST API both run OmniVoice TTS inference, but differ in execution environment, batch handling, parallelism model, and result delivery.**

The VoiceStudio project (debpalash/VoiceStudio) provides two interfaces for running OmniVoice text-to-speech: a command-line interface for local execution and a FastAPI-based REST API for remote access. While both leverage the same core generation logic in `OmniVoice.generate`, they serve different operational needs and scale differently. Understanding these differences helps you choose the right approach for your inference workload.

## How the CLI and REST API Are Invoked

### CLI Entry Points

The VoiceStudio CLI provides two distinct commands:

1. **`omnivoice-infer`** — single-sample inference, implemented in [`omnivoice/cli/infer.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer.py)
2. **`omnivoice-infer-batch`** — batch processing, implemented in [`omnivoice/cli/infer_batch.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer_batch.py)

Both execute locally on the machine where invoked. The binary loads the model directly from a checkpoint path or Hugging Face repository and runs inference in-process.

### REST API Entry Point

The REST API runs inside a FastAPI server started with `uvicorn`. The server loads the model once at startup and serves requests over HTTP. The primary endpoint is `/v1/audio/speech`, defined in [`omnivoice/api/v1.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/v1.py).

```bash

# Start the server

uvicorn omnivoice.api.server:app --host 0.0.0.0 --port 3900

```

The server entry point resides in [`omnivoice/api/server.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/server.py), which bootstraps the FastAPI application.

## Batch Handling Comparison

### CLI Batch Processing

The batch CLI reads a **JSON-L file** (`--test_list`) containing multiple samples. It clusters inputs by duration or fixed batch size, then dispatches them to a pool of worker processes using `ProcessPoolExecutor`.

```bash

# Prepare test list

cat > test.jsonl <<EOF
{"id":"sample1","text":"Hello world!","ref_audio":"ref1.wav","ref_text":"Hello world!"}
{"id":"sample2","text":"How are you?","ref_audio":"ref2.wav","ref_text":"How are you?"}
EOF

# Run batch inference

omnivoice-infer-batch \
  --model k2-fsa/OmniVoice \
  --test_list test.jsonl \
  --res_dir results/ \
  --batch_duration 1000.0 \
  --nj_per_gpu 2

```

The [`omnivoice/cli/infer_batch.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer_batch.py) script handles argument parsing, worker pool creation, sample clustering, and WAV file output to `results/`.

### REST API Batch Processing

The REST API accepts a JSON payload with an array of `input` objects. The server groups requests internally and returns audio data for each entry.

```bash
curl -X POST http://localhost:3900/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
        "model":"k2-fsa/OmniVoice",
        "batch":[
          {"id":"sample1","text":"Hello world!","ref_audio":"/path/ref1.wav","ref_text":"Hello world!"},
          {"id":"sample2","text":"How are you?","ref_audio":"/path/ref2.wav","ref_text":"How are you?"}
        ],
        "guidance_scale":2.0,
        "num_step":32
      }' \
  -o response.json

```

The handler in [`omnivoice/api/v1.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/v1.py) extracts the batch list, calls `OmniVoice.generate`, and returns either base64-encoded WAV data or URLs to stored files.

## Parallelism Model

### CLI Parallelism

The CLI implements **explicit multi-process parallelism** through these flags:

- `--nj_per_gpu` — worker processes per GPU
- `--batch_duration` — target total duration per batch cluster
- `--batch_size` — fixed sample count per batch

This design exploits multiple GPUs on the same node directly, with each worker maintaining its own model instance.

### REST API Parallelism

Parallelism in the REST API is managed by:

- Uvicorn worker processes (server-level)
- Device mapping within the model
- Request queue handling by FastAPI

Scaling typically involves launching multiple server workers rather than controlling GPU allocation per request.

## Result Delivery

| Aspect | CLI (`omnivoice-infer-batch`) | REST API |
|--------|------------------------------|----------|
| Output format | Individual `.wav` files | Base64-encoded audio or file URLs |
| Destination | Local directory (`--res_dir`) | HTTP response body |
| Additional output | Console summary with RTF (real-time factor) | JSON response with per-sample results |
| Streaming | Not supported | Supported if client requests it |

After CLI execution, you'll find files like `results/sample1.wav` and `results/sample2.wav`, plus timing statistics printed to stdout. The API returns everything in a single HTTP response.

## Configuration Parameters

Both interfaces accept the same generation parameters with different syntax:

**CLI flags:**

```bash
omnivoice-infer-batch \
  --guidance_scale 2.0 \
  --num_step 32 \
  --preprocess_prompt true

```

**REST API JSON fields:**

```json
{
  "guidance_scale": 2.0,
  "num_step": 32,
  "preprocess_prompt": true
}

```

The parameter names and valid ranges are identical; only the transmission mechanism differs.

## Error Handling

### CLI Error Behavior

Exceptions are caught per-sample in [`omnivoice/cli/infer_batch.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer_batch.py). Failed samples are logged and the batch continues. Check console output or redirected logs for failure details.

### REST API Error Behavior

HTTP status codes indicate outcomes:
- `4xx` — client errors (invalid parameters, malformed JSON)
- `5xx` — server errors (inference failures, model loading issues)

The response body contains structured error messages for programmatic handling.

## When to Use Each Interface

**Choose the CLI when:**
- Processing large offline corpora
- Running scriptable pipelines
- Direct GPU access is available
- You need maximum throughput via `ProcessPoolExecutor`
- Output files must be organized in local directories

**Choose the REST API when:**
- Integrating with web or mobile applications
- Running inference as a cloud service
- Clients need programmatic access without local model installation
- OpenAI-compatible client libraries are preferred
- Request distribution across network infrastructure is required

## Core Implementation Files

| Component | Source File | Purpose |
|-----------|-------------|---------|
| Single-sample CLI | [`omnivoice/cli/infer.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer.py) | One-shot inference with argument parsing |
| Batch CLI | [`omnivoice/cli/infer_batch.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer_batch.py) | JSON-L processing, worker pools, file output |
| API router | [`omnivoice/api/v1.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/v1.py) | FastAPI routes for `/v1/audio/speech` |
| Server bootstrap | [`omnivoice/api/server.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/server.py) | Uvicorn application entry point |
| Model definition | [`omnivoice/models/omnivoice.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/omnivoice.py) | `OmniVoice.from_pretrained` and `generate` method |
| Audio utilities | [`omnivoice/utils/audio.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/utils/audio.py) | `load_audio`, `save_audio` shared by both interfaces |

These files demonstrate that `OmniVoice.generate` in [`omnivoice/models/omnivoice.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/omnivoice.py) serves as the unified core, while interface-specific code handles I/O and orchestration.

## Summary

- **Execution**: CLI runs locally in-process; REST API serves over HTTP via FastAPI
- **Batch input**: CLI uses JSON-L files; REST API accepts JSON payloads
- **Parallelism**: CLI uses explicit `ProcessPoolExecutor` with GPU workers; REST API relies on server workers
- **Results**: CLI writes `.wav` files to disk; REST API returns audio in HTTP responses
- **Shared core**: Both use `OmniVoice.generate` from [`omnivoice/models/omnivoice.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/omnivoice.py)
- **Source locations**: CLI in `omnivoice/cli/infer*.py`, API in [`omnivoice/api/v1.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/api/v1.py)

## Frequently Asked Questions

### Can I use the REST API for the same large-scale batch jobs as the CLI?

Not optimally. The REST API handles batch requests but lacks the CLI's explicit multi-GPU process pooling. For processing thousands of samples offline, the CLI's `omnivoice-infer-batch` with `--nj_per_gpu` and `--batch_duration` controls provides better throughput and resource utilization. Use the API when network accessibility matters more than raw throughput.

### Do the CLI and REST API produce identical audio output?

Yes, when given identical parameters. Both call the same `OmniVoice.generate` method in [`omnivoice/models/omnivoice.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/models/omnivoice.py). Any differences stem from parameter values, not the inference path. Verify by running the same sample through both interfaces with matching `guidance_scale`, `num_step`, and `preprocess_prompt` settings.

### How do I choose between single-sample and batch CLI commands?

Use `omnivoice-infer` (from [`omnivoice/cli/infer.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer.py)) for one-off tests or interactive workflows. Use `omnivoice-infer-batch` (from [`omnivoice/cli/infer_batch.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice/cli/infer_batch.py)) when processing multiple samples, as it implements efficient clustering and parallel execution across GPUs that single-sample invocation cannot match.