How to Select the Compute Type for ASR in VoiceStudio: int8, float16, and float32 Explained

Set the OMNIVOICE_ASR_COMPUTE_TYPE environment variable to float16, int8_float16, int8, or float32 to manually select the ASR compute type, or let VoiceStudio automatically choose based on your GPU and available VRAM.

VoiceStudio intelligently manages numeric precision for automatic speech recognition (ASR) by selecting optimal compute types based on hardware capabilities and memory constraints. The selection logic implemented in backend/services/asr_backend.py automatically balances transcription speed and accuracy across CUDA GPUs and CPU environments. Understanding how to select the compute type for ASR allows you to override defaults for specific performance requirements or memory limitations.

How VoiceStudio Automatically Selects the ASR Compute Type

VoiceStudio determines the optimal compute type through device detection and VRAM analysis before loading models. The ASRBackend._candidate_compute_types() method (line 257) defines the priority chains for different hardware configurations.

CUDA GPU Defaults

On CUDA-capable GPUs, VoiceStudio attempts the following precision chain:

  1. float16 – 16-bit floating point (fastest when hardware supports native half-precision)
  2. int8_float16 – 8-bit integer weights with float16 activations (GPU-only hybrid)
  3. int8 – Pure 8-bit integer inference (maximum memory efficiency)

The engine attempts each type in sequence, falling back to the next option if the GPU lacks efficient support for higher precision modes.

CPU Defaults

For CPU-only environments or non-CUDA devices, the backend follows a conservative chain:

  1. int8 – 8-bit quantized inference (recommended for CPU)
  2. float32 – 32-bit floating point (fallback for compatibility)

Manual Override with Environment Variables

You can explicitly force a specific compute type by setting the OMNIVOICE_ASR_COMPUTE_TYPE environment variable before initializing the backend. When this variable is detected, VoiceStudio bypasses automatic detection and uses your specified precision, provided the hardware supports it.

Valid values include:

  • float16 – Full half-precision floating point (requires modern NVIDIA GPU)
  • int8_float16 – Mixed quantization (GPU-only)
  • int8 – Integer quantization (works on GPU and CPU)
  • float32 – Full single-precision floating point (CPU fallback)

The variable is read early during backend initialization via the _pick_device() helper. If the requested compute type is incompatible with the available hardware, VoiceStudio logs a warning and falls back to the next viable option (lines 728–733).

import os

# Force int8 inference for consistent memory usage

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "int8"

# Initialize the backend (will use int8 regardless of GPU capability)

from backend.services.asr_backend import get_asr_backend
asr = get_asr_backend()

VRAM-Budget Protection and Automatic Degradation

VoiceStudio implements a VRAM-guard mechanism that prevents out-of-memory errors by degrading compute types when necessary. Before loading a model, the backend compares available VRAM against the _CUDA_VRAM_BUDGET_GB table (line 728).

If the selected compute type would exceed the memory budget, the engine automatically steps down the precision chain:


float16 → int8_float16 → int8

This automatic degradation ensures transcription jobs complete successfully even on GPUs with limited memory, trading marginal accuracy for system stability.

Fallback Behavior for Unsupported Hardware

When a compute type is unsupported by the underlying hardware (for example, requesting float16 on a GPU without efficient half-precision support), VoiceStudio triggers a structured fallback sequence.

The load_asr routine catches hardware incompatibility errors and consults the error mapping in backend/core/failure.py (line 90). The system logs a COMPUTE_TYPE_UNSUPPORTED message and automatically retries with the next lower-precision option in the candidate chain until finding a compatible configuration.

Practical Code Examples

Allow VoiceStudio to detect your hardware and select the optimal compute type:

import os

# Ensure no override is set

os.environ.pop("OMNIVOICE_ASR_COMPUTE_TYPE", None)

from backend.services.asr_backend import get_asr_backend
asr = get_asr_backend()  # Will select float16/int8 based on GPU capability

Force float16 on Compatible GPUs

Explicitly select half-precision for maximum speed on modern NVIDIA hardware:

import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "float16"
asr = get_asr_backend()  # Loads with torch.float16 dtype

CPU-Only float32 Mode

Force 32-bit precision for CPU inference or maximum numerical accuracy:

import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "float32"
asr = get_asr_backend()  # Loads in float32 mode

Mixed-Mode GPU Quantization

Use int8_float16 for balanced performance and memory usage:

import os

os.environ["OMNIVOICE_ASR_COMPUTE_TYPE"] = "int8_float16"
asr = get_asr_backend()  # 8-bit weights, 16-bit activations

Summary

  • Automatic selection chooses between float16, int8_float16, and int8 on GPUs, or int8 and float32 on CPUs, based on hardware capabilities.
  • Manual override via OMNIVOICE_ASR_COMPUTE_TYPE forces a specific precision when you need consistent behavior across different environments.
  • VRAM protection automatically degrades compute types when memory constraints are detected, preventing out-of-memory crashes.
  • Hardware fallback sequences ensure ASR models load successfully even when the requested precision is unsupported.

Frequently Asked Questions

What is the fastest compute type for ASR transcription in VoiceStudio?

float16 provides the fastest inference on modern NVIDIA GPUs with native half-precision support (Tensor Cores). However, if your GPU lacks efficient float16 support or has limited VRAM, int8_float16 or pure int8 may offer better throughput due to reduced memory bandwidth requirements.

Can I use float16 on CPU-only systems?

No, float16 and int8_float16 are GPU-only compute types in VoiceStudio. If you attempt to set OMNIVOICE_ASR_COMPUTE_TYPE=float16 on a CPU-only machine, the backend will fall back to float32 or int8 depending on the failure handling logic in backend/core/failure.py.

Why does VoiceStudio ignore my compute type setting sometimes?

If the requested compute type exceeds available VRAM or is unsupported by your specific GPU architecture, VoiceStudio automatically degrades to a compatible alternative. Check the logs for COMPUTE_TYPE_UNSUPPORTED warnings around the load_asr initialization to see if automatic fallback occurred.

How do I verify which compute type is currently active?

Inspect the backend logs during model initialization. The ASRBackend logs the selected dtype (e.g., torch.float16, torch.int8) when loading models through backend/services/model_manager.py. You can also check the runtime environment variables to confirm OMNIVOICE_ASR_COMPUTE_TYPE was set before import.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →