# How to Select the Compute Type for ASR in VoiceStudio: int8, float16, and float32 Explained

> Learn to select ASR compute types like int8, float16, and float32 in VoiceStudio. Optimize performance by manually setting the OMNIVOICE_ASR_COMPUTE_TYPE variable or letting VoiceStudio auto-detect.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**Set the `OMNIVOICE_ASR_COMPUTE_TYPE` environment variable to `float16`, `int8_float16`, `int8`, or `float32` to manually select the ASR compute type, or let VoiceStudio automatically choose based on your GPU and available VRAM.**

VoiceStudio intelligently manages numeric precision for automatic speech recognition (ASR) by selecting optimal compute types based on hardware capabilities and memory constraints. The selection logic implemented in [`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py) automatically balances transcription speed and accuracy across CUDA GPUs and CPU environments. Understanding how to select the compute type for ASR allows you to override defaults for specific performance requirements or memory limitations.

## How VoiceStudio Automatically Selects the ASR Compute Type

VoiceStudio determines the optimal compute type through device detection and VRAM analysis before loading models. The `ASRBackend._candidate_compute_types()` method (line [257](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py#L257)) defines the priority chains for different hardware configurations.

### CUDA GPU Defaults

On CUDA-capable GPUs, VoiceStudio attempts the following precision chain:

1. **float16** – 16-bit floating point (fastest when hardware supports native half-precision)
2. **int8_float16** – 8-bit integer weights with float16 activations (GPU-only hybrid)
3. **int8** – Pure 8-bit integer inference (maximum memory efficiency)

The engine attempts each type in sequence, falling back to the next option if the GPU lacks efficient support for higher precision modes.

### CPU Defaults

For CPU-only environments or non-CUDA devices, the backend follows a conservative chain:

1. **int8** – 8-bit quantized inference (recommended for CPU)
2. **float32** – 32-bit floating point (fallback for compatibility)

## Manual Override with Environment Variables

You can explicitly force a specific compute type by setting the **`OMNIVOICE_ASR_COMPUTE_TYPE`** environment variable before initializing the backend. When this variable is detected, VoiceStudio bypasses automatic detection and uses your specified precision, provided the hardware supports it.

Valid values include:

- `float16` – Full half-precision floating point (requires modern NVIDIA GPU)
- `int8_float16` – Mixed quantization (GPU-only)
- `int8` – Integer quantization (works on GPU and CPU)
- `float32` – Full single-precision floating point (CPU fallback)

The variable is read early during backend initialization via the `_pick_device()` helper. If the requested compute type is incompatible with the available hardware, VoiceStudio logs a warning and falls back to the next viable option (lines [728–733](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py#L728)).

```python
import os

# Force int8 inference for consistent memory usage

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "int8"

# Initialize the backend (will use int8 regardless of GPU capability)

from backend.services.asr_backend import get_asr_backend
asr = get_asr_backend()

```

## VRAM-Budget Protection and Automatic Degradation

VoiceStudio implements a VRAM-guard mechanism that prevents out-of-memory errors by degrading compute types when necessary. Before loading a model, the backend compares available VRAM against the `_CUDA_VRAM_BUDGET_GB` table (line [728](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py#L728)).

If the selected compute type would exceed the memory budget, the engine automatically steps down the precision chain:

```

float16 → int8_float16 → int8

```

This automatic degradation ensures transcription jobs complete successfully even on GPUs with limited memory, trading marginal accuracy for system stability.

## Fallback Behavior for Unsupported Hardware

When a compute type is unsupported by the underlying hardware (for example, requesting `float16` on a GPU without efficient half-precision support), VoiceStudio triggers a structured fallback sequence.

The `load_asr` routine catches hardware incompatibility errors and consults the error mapping in [`backend/core/failure.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/failure.py) (line [90](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/failure.py#L90)). The system logs a `COMPUTE_TYPE_UNSUPPORTED` message and automatically retries with the next lower-precision option in the candidate chain until finding a compatible configuration.

## Practical Code Examples

### Automatic Selection (Recommended)

Allow VoiceStudio to detect your hardware and select the optimal compute type:

```python
import os

# Ensure no override is set

os.environ.pop("OMNIVOICE_ASR_COMPUTE_TYPE", None)

from backend.services.asr_backend import get_asr_backend
asr = get_asr_backend()  # Will select float16/int8 based on GPU capability

```

### Force float16 on Compatible GPUs

Explicitly select half-precision for maximum speed on modern NVIDIA hardware:

```python
import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "float16"
asr = get_asr_backend()  # Loads with torch.float16 dtype

```

### CPU-Only float32 Mode

Force 32-bit precision for CPU inference or maximum numerical accuracy:

```python
import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "float32"
asr = get_asr_backend()  # Loads in float32 mode

```

### Mixed-Mode GPU Quantization

Use `int8_float16` for balanced performance and memory usage:

```python
import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "int8_float16"
asr = get_asr_backend()  # 8-bit weights, 16-bit activations

```

## Summary

- **Automatic selection** chooses between `float16`, `int8_float16`, and `int8` on GPUs, or `int8` and `float32` on CPUs, based on hardware capabilities.
- **Manual override** via `OMNIVOICE_ASR_COMPUTE_TYPE` forces a specific precision when you need consistent behavior across different environments.
- **VRAM protection** automatically degrades compute types when memory constraints are detected, preventing out-of-memory crashes.
- **Hardware fallback** sequences ensure ASR models load successfully even when the requested precision is unsupported.

## Frequently Asked Questions

### What is the fastest compute type for ASR transcription in VoiceStudio?

**float16** provides the fastest inference on modern NVIDIA GPUs with native half-precision support (Tensor Cores). However, if your GPU lacks efficient float16 support or has limited VRAM, **int8_float16** or pure **int8** may offer better throughput due to reduced memory bandwidth requirements.

### Can I use float16 on CPU-only systems?

No, `float16` and `int8_float16` are GPU-only compute types in VoiceStudio. If you attempt to set `OMNIVOICE_ASR_COMPUTE_TYPE=float16` on a CPU-only machine, the backend will fall back to `float32` or `int8` depending on the failure handling logic in [`backend/core/failure.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/failure.py).

### Why does VoiceStudio ignore my compute type setting sometimes?

If the requested compute type exceeds available VRAM or is unsupported by your specific GPU architecture, VoiceStudio automatically degrades to a compatible alternative. Check the logs for `COMPUTE_TYPE_UNSUPPORTED` warnings around the `load_asr` initialization to see if automatic fallback occurred.

### How do I verify which compute type is currently active?

Inspect the backend logs during model initialization. The `ASRBackend` logs the selected dtype (e.g., `torch.float16`, `torch.int8`) when loading models through [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py). You can also check the runtime environment variables to confirm `OMNIVOICE_ASR_COMPUTE_TYPE` was set before import.