# Voicebox Generation Queue and Task Scheduling: Serial GPU Execution Explained

> Understand Voicebox generation queue and task scheduling. Learn how serial GPU execution prevents resource contention for efficient text-to-speech.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: internals
- Published: 2026-04-14

---

**Voicebox uses a single global `asyncio.Queue` processed by a dedicated worker coroutine to ensure only one text-to-speech generation runs at a time, preventing GPU resource contention.**

Voicebox (jamiepine/voicebox) is an open-source text-to-speech platform that performs inference on GPU hardware. Because running multiple TTS jobs simultaneously causes resource conflicts, the application implements a strict serial execution model through an in-process generation queue and background worker that processes one coroutine at a time.

## Queue Architecture and Worker Initialization

When the FastAPI application starts, it initializes a global queue and background worker to manage all generation tasks. This architecture guarantees that GPU-intensive inference operations never overlap.

### Global Queue Creation

In [`backend/app.py`](https://github.com/jamiepine/voicebox/blob/main/backend/app.py) lines 99–100, the `startup_event` handler calls `init_queue()` to create the scheduling infrastructure:

```python

# From backend/app.py

@app.on_event("startup")
async def startup_event():
    init_queue()  # Creates global asyncio.Queue

```

This function instantiates a module-level `asyncio.Queue` object and launches the background worker task that will consume from it for the entire application lifetime.

### The Serial Worker Loop

The core execution logic resides in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) lines 24–34 inside the `_generation_worker` coroutine:

```python

# Conceptual implementation from backend/services/task_queue.py

async def _generation_worker():
    while True:
        coro = await _generation_queue.get()
        try:
            await coro
        except Exception as e:
            logger.error(f"Generation failed: {e}")
        finally:
            _generation_queue.task_done()

```

Because this worker runs as a single `asyncio` task, it pulls one coroutine from the queue, awaits its completion, and only then retrieves the next. This sequential processing ensures exclusive GPU access for each generation job.

## Enqueuing and Executing Generations

Client requests do not execute inference directly; instead they package the work as coroutines and submit them to the global queue.

### The Public API Surface

The `enqueue_generation(coro)` helper in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) lines 36–38 provides the safe entry point:

```python

# From backend/services/task_queue.py

def enqueue_generation(coro):
    _generation_queue.put_nowait(coro)

```

All components that need TTS processing call this function rather than awaiting `run_generation` directly. The `put_nowait` method places the coroutine into the queue immediately without blocking the caller.

### The Generation Pipeline

The heavy lifting occurs in `run_generation()` located in [`backend/services/generation.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/generation.py) lines 28–45. This function:

- Loads the specified TTS model into GPU memory
- Builds voice prompts from profile data
- Performs chunked audio synthesis
- Applies optional normalization and audio effects
- Persists the final audio file and updates the database

As noted in the docstring at line 48, this function is **designed to be enqueued** and should only be awaited inside the queue worker to maintain serial execution.

## API Integration and Task Tracking

The REST layer bridges HTTP requests to the queue system while providing real-time status monitoring through a task tracking utility.

### REST Endpoint Flow

When clients POST to the generation endpoint, [`backend/routes/generations.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/generations.py) lines 84–100 orchestrates the workflow:

1. Generates a new UUID for the generation
2. Creates a database record with pending status
3. Starts tracking via `TaskManager.start_generation()`
4. Enqueues the work: `enqueue_generation(run_generation(...))`

This pattern applies to standard generation requests, retries, and regeneration endpoints. The API returns immediately with the generation ID while the actual inference runs asynchronously in the queue.

### TaskManager Lifecycle Management

The `TaskManager` class in [`backend/utils/tasks.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/tasks.py) maintains dictionaries of active downloads and generations. When `run_generation` begins, it marks the task active; when it completes or fails, the `finally` block in [`backend/services/generation.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/generation.py) lines 136–139 calls `task_manager.complete_generation(generation_id)` to clear the entry:

```python

# From backend/services/generation.py

try:
    # ... inference logic ...

    pass
except Exception as e:
    task_manager.fail_generation(generation_id, str(e))
    raise
finally:
    task_manager.complete_generation(generation_id)  # Lines 136-139

```

This allows the frontend to query `/tasks` endpoints for live progress updates while the background worker handles the actual GPU work.

## Implementation Examples

### Enqueue a Generation from Python

Call the public helper to schedule work from scripts or other coroutines:

```python
from voicebox.backend.services.task_queue import enqueue_generation
from voicebox.backend.services.generation import run_generation

# Build the coroutine with desired parameters

coro = run_generation(
    generation_id="c1a2b3",
    profile_id="profile-42",
    text="Hello, world!",
    language="en",
    engine="qwen",
    model_size="1.7B",
    seed=None,
    normalize=True,
    effects_chain=None,
    instruct=None,
    mode="generate",
)

# Push it onto the serial queue

enqueue_generation(coro)

```

### Trigger via HTTP API

Submit a generation request through the REST endpoint:

```bash
curl -X POST http://localhost:8000/generate \
  -H "Content-Type: application/json" \
  -d '{
        "profile_id":"profile-42",
        "text":"Hello, world!",
        "language":"en",
        "engine":"qwen",
        "model_size":"1.7B",
        "normalize":true
      }'

```

The endpoint stores the database record, initializes task tracking, and internally calls `enqueue_generation(run_generation(...))`, ensuring the TTS inference runs exclusively on the GPU.

## Summary

- **Single-threaded queue**: Voicebox uses a global `asyncio.Queue` with one dedicated worker to enforce serial GPU access
- **Separation of concerns**: API endpoints enqueue coroutines immediately but return responses; actual inference runs asynchronously in the background
- **Automatic lifecycle**: `TaskManager` tracks active generations from enqueue time through completion, with cleanup handled in `finally` blocks
- **Non-blocking interface**: `enqueue_generation()` uses `put_nowait()` to accept jobs without blocking the event loop

## Frequently Asked Questions

### Why does Voicebox process generations serially instead of concurrently?

Voicebox performs text-to-speech inference on GPU hardware. Running multiple models or batches simultaneously causes memory contention and out-of-memory errors. The serial queue in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) guarantees that only one generation accesses the GPU at a time, ensuring predictable resource usage and preventing job failures.

### How can I check the status of a queued generation?

The application tracks all active generations through the `TaskManager` class in [`backend/utils/tasks.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/tasks.py). When you enqueue a job via the REST API at [`backend/routes/generations.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/generations.py), the endpoint returns a generation ID. Query the `/tasks` endpoints to retrieve real-time status updates, as the worker updates the task state when `run_generation` completes or fails.

### What happens if a generation fails inside the queue?

The `_generation_worker` in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) wraps each coroutine execution in a try-except block that logs exceptions. Additionally, `run_generation` in [`backend/services/generation.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/generation.py) uses a `finally` block (lines 136–139) to call `task_manager.complete_generation()`, ensuring that failed jobs clear their active status and release tracking resources regardless of success or failure.

### Is the generation queue persistent across application restarts?

No. The queue is an in-memory `asyncio.Queue` created by `init_queue()` during the FastAPI startup event in [`backend/app.py`](https://github.com/jamiepine/voicebox/blob/main/backend/app.py) lines 99–100. When the application stops, the queue and any unprocessed coroutines are lost. For persistence, the system relies on the database records created before enqueueing, which allow regeneration or retry from the stored state.