# How to Implement Error Handling and Fallback for Supertonic TTS Inference Failures

> Implement robust error handling and fallback for Supertonic TTS inference. Use try-catch, multi-tier fallbacks, and optional error states to prevent hangs.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Wrap the `_infer` method in try-catch blocks, implement a three-tier fallback strategy (retry with reduced steps, switch to CPU provider, then return silence), and return result tuples containing optional error states to prevent application hangs during ONNX runtime failures.**

Supertonic is a high-performance text-to-speech engine built on ONNX Runtime that powers real-time voice synthesis across Python, Go, Java, and Swift applications. When deploying **Supertonic TTS inference** in production, unhandled exceptions from missing model files, GPU memory exhaustion, or input shape mismatches can crash your application or freeze the user interface. Implementing robust error handling and fallback mechanisms ensures your service remains resilient even when the underlying neural network sessions fail.

## Understanding the Core Inference Architecture

Supertonic's TTS engine orchestrates four distinct ONNX inference sessions: the **duration predictor**, **text encoder**, **vector estimator**, and **vocoder**. In the Python SDK, the `_infer` method in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (lines 177-186) coordinates these sessions, while equivalent implementations exist in [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) (lines 678-702), [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java) (lines 813-825), and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) (lines 692-715). Each session's `run` call can throw `OrtError` exceptions when model files are corrupted, input tensors have mismatched shapes, or the GPU runs out of memory.

## Guarding the Inference Pipeline with Exception Handling

To prevent cascading failures, wrap every ONNX session call within the `_infer` method in language-specific exception handlers. The goal is to catch errors at the session level before they propagate to the top-level API.

### Python SDK Implementation

In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the `_infer` method executes the vocoder and other sessions without fallback protection. Create a wrapper method that catches exceptions and implements tiered recovery:

```python
def infer_with_fallback(
    self,
    text_list,
    lang_list,
    style,
    total_step,
    speed=1.05,
    fallback_steps: int = 1,
    fallback_audio: np.ndarray | None = None,
) -> tuple[np.ndarray, np.ndarray, Exception | None]:
    """
    Runs inference and falls back on failure.
    """
    try:
        return self._infer(text_list, lang_list, style, total_step, speed) + (None,)
    except Exception as primary_err:
        print(f"[Supertonic] Primary inference error: {primary_err}")
        
        # Retry with reduced step count

        try:
            wav, dur = self._infer(
                text_list, lang_list, style, fallback_steps, speed
            )
            return wav, dur, None
        except Exception as retry_err:
            print(f"[Supertonic] Fallback inference error: {retry_err}")
            
            # Return silence as final fallback

            if fallback_audio is None:
                silence_len = int(0.5 * self.sample_rate)
                fallback_audio = np.zeros((1, silence_len), dtype=np.float32)
            return fallback_audio, np.array([0.5]), retry_err

```

This pattern returns a three-element tuple containing the audio data, duration, and an optional error object, allowing the caller to determine whether to play the synthesized speech or handle the failure.

### Go SDK Implementation

For Go applications, modify the `_infer` method in [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) to return errors explicitly rather than panicking:

```go
func (tts *TextToSpeech) InferWithFallback(
    textList []string,
    langList []string,
    style *Style,
    totalStep int,
    fallbackStep int,
) ([]float32, []float32, error) {
    wav, dur, err := tts._infer(textList, langList, style, totalStep, 1.05)
    if err == nil {
        return wav, dur, nil
    }
    
    fmt.Printf("[Supertonic] Primary inference error: %v\n", err)
    
    // Retry with reduced steps
    wav, dur, retryErr := tts._infer(textList, langList, style, fallbackStep, 1.05)
    if retryErr == nil {
        return wav, dur, nil
    }
    
    // Final fallback - silence buffer
    fmt.Printf("[Supertonic] Fallback inference error: %v\n", retryErr)
    silence := make([]float32, int(0.5*float32(tts.SampleRate)))
    return silence, []float32{0.5}, retryErr
}

```

### Java SDK Pattern

In [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java), wrap the `infer` method in try-catch blocks to intercept runtime exceptions:

```java
public TTSResult inferWithFallback(..., int fallbackSteps) {
    try {
        return infer(...);
    } catch (Exception e) {
        System.err.println("[Supertonic] Primary inference error: " + e);
        try {
            return infer(..., fallbackSteps);
        } catch (Exception retry) {
            System.err.println("[Supertonic] Fallback inference error: " + retry);
            float[] silence = new float[(int) (0.5f * sampleRate)];
            return new TTSResult(silence, 0.5f, retry);
        }
    }
}

```

### Swift SDK Pattern

For iOS and macOS applications using [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), use `do-catch` syntax:

```swift
func inferWithFallback(..., fallbackSteps: Int = 1) -> (wav: [Float], duration: Float, error: Error?) {
    do {
        return try infer(...)
    } catch {
        print("[Supertonic] Primary inference error: \(error)")
        do {
            return try infer(..., totalStep: fallbackSteps)
        } catch let fallbackError {
            print("[Supertonic] Fallback inference error: \(fallbackError)")
            let silence = [Float](repeating: 0, count: Int(0.5 * sampleRate))
            return (silence, 0.5, fallbackError)
        }
    }
}

```

## Tiered Fallback Strategies for Production

Implementing multiple recovery layers ensures graceful degradation when **Supertonic TTS inference failures** occur.

### Retry with Reduced Step Count

When GPU memory limits cause failures, retry with `total_step = max(1, total_step // 2)`. This trades synthesis quality for memory efficiency and often succeeds when the full step count fails due to memory constraints in the vocoder session.

### CPU Provider Switching

When GPU execution throws `OrtError` during session creation in `load_onnx_all`, catch the exception and recreate the ONNX Runtime environment with CPU-only providers. This avoids GPU memory constraints entirely at the cost of inference speed, making it ideal for critical paths where any audio output is preferable to a crash.

### Static Audio Fallback

As a last resort, return a pre-generated silence buffer (0.5 seconds of zeros) or a cached "unable to synthesize" audio clip. This prevents the client application from hanging and allows the UI to display an error message while maintaining audio stream continuity.

## Summary

- **Wrap all ONNX calls** in the `_infer` method with language-specific exception handlers to catch `OrtError` and memory failures.
- **Implement tiered fallbacks**: first retry with reduced `total_step`, then switch to CPU providers, then return static silence.
- **Return error tuples** rather than throwing exceptions to the caller, allowing the application to decide how to handle synthesis failures.
- **Log failures** with model names and step counts to identify systematic issues with specific model bundles or hardware configurations.

## Frequently Asked Questions

### What causes Supertonic TTS inference failures?

Failures typically stem from three sources: missing or corrupted ONNX model files during `load_onnx_all`, GPU out-of-memory errors when running the vocoder session, and input shape mismatches between the text encoder and duration predictor. According to the source code in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), these manifest as exceptions during the `run` calls of the four inference sessions.

### How does reducing total_step help with inference errors?

Reducing `total_step` decreases the computational complexity of the diffusion process in the vocoder, lowering both memory consumption and execution time. As implemented in the fallback wrappers, setting `total_step` to 1 or 2 often succeeds on constrained hardware where the default 10+ steps would exhaust available GPU memory.

### Can I fall back to a different ONNX provider instead of CPU?

Yes. If the CUDA provider fails during session initialization in `load_onnx_all`, you can catch the exception and recreate the session with the `CPUExecutionProvider` (Python) or equivalent CPU providers in other languages. This is often more reliable than retrying with reduced steps, though it significantly impacts real-time performance.

### How should I handle errors in the top-level TTS API?

Always return result objects or tuples containing both the audio data and an optional error field. In Python, use `(audio, duration, error)`; in Go, return `([]float32, []float32, error)`; in Java and Swift, return custom result objects with `error` properties. This pattern allows higher-level UI code to detect failures and display appropriate user feedback while playing fallback audio if available.