Voicebox Async Generation with SSE Streaming: Implementation Guide
Voicebox implements asynchronous text-to-speech generation using Server-Sent Events (SSE) to stream real-time progress updates from the FastAPI backend to the React frontend, enabling fire-and-forget generation with automatic audio playback upon completion.
Voicebox by jamiepine delivers a fully asynchronous text-to-speech pipeline that eliminates polling overhead through persistent SSE connections. When users initiate generation, the backend spawns background TTS workers while the frontend subscribes to event streams that report granular progress states including loading_model, generating, completed, failed, and not_found. This architecture enables real-time UI updates without page reloads and supports automatic audio playback the moment synthesis finishes.
Architecture Overview
The async generation system spans both frontend and backend layers, with clean separation between connection management, business logic, and data persistence.
| Component | Responsibility | Primary Files |
|---|---|---|
| React UI | Initiates generation, manages pending IDs, consumes SSE streams, handles auto-playback logic | app/src/lib/hooks/useGenerationProgress.ts |
| API Client | Constructs SSE endpoint URLs and audio download paths | app/src/lib/api/client.ts |
| FastAPI Backend | Accepts requests, schedules background jobs, exposes SSE status streams and direct audio streaming | backend/routes/generations.py |
| Task Workers | Execute heavy-weight TTS computation via generate_chunked and update database records |
backend/utils/chunked_tts.py |
| Database | Persists DBGeneration rows with status, duration, and error fields |
backend/database/models.py |
The SSE stream maintains a lightweight, unidirectional HTTP connection that pushes JSON payloads whenever generation status changes. The UI tracks active connections using a Map<string, EventSource> keyed by generation ID, ensuring resources are released promptly when jobs complete or fail.
SSE Data Flow
The generation lifecycle follows a strict event-driven protocol:
- Generation Request: The UI POSTs a
GenerationRequestto/generate. The server creates aDBGenerationrow with statusloading_modeland returns the UUID. - Pending Tracking: The client adds the ID to a Zustand store (
useGenerationStore.pendingGenerationIds). - SSE Connection: The
useGenerationProgresshook opens anEventSourcetoGET /generate/{id}/status. - Server Loop: The
event_streamcoroutine inbackend/routes/generations.pyqueries the database every second, yields JSON payloads, and sleeps viaawait asyncio.sleep(1)until status reachescompletedorfailed. - Completion Handling: Upon receiving the final event, the hook closes the connection, invalidates React Query caches, optionally adds the audio to stories, and triggers auto-playback if enabled.
This fire-and-forget pattern ensures the UI never blocks waiting for generation while providing real-time feedback.
Client-Side Implementation
The useGenerationProgress Hook
The core of the frontend implementation resides in app/src/lib/hooks/useGenerationProgress.ts. This hook manages the lifecycle of SSE connections and coordinates state updates across the application.
// app/src/lib/hooks/useGenerationProgress.ts
export function useGenerationProgress() {
const queryClient = useQueryClient();
const { toast } = useToast();
const pendingIds = useGenerationStore(s => s.pendingGenerationIds);
const removePendingGeneration = useGenerationStore(s => s.removePendingGeneration);
const removePendingStoryAdd = useGenerationStore(s => s.removePendingStoryAdd);
const isPlaying = usePlayerStore(s => s.isPlaying);
const setAudioWithAutoPlay = usePlayerStore(s => s.setAudioWithAutoPlay);
const autoplayOnGenerate = useServerStore(s => s.autoplayOnGenerate);
const isPlayingRef = useRef(isPlaying);
const autoplayRef = useRef(autoplayOnGenerate);
isPlayingRef.current = isPlaying;
autoplayRef.current = autoplayOnGenerate;
const eventSourcesRef = useRef<Map<string, EventSource>>(new Map());
// Cleanup all connections on unmount
useEffect(() => {
const sources = eventSourcesRef.current;
return () => {
for (const source of sources.values()) source.close();
sources.clear();
};
}, []);
// Manage SSE connections based on pending IDs
useEffect(() => {
const currentSources = eventSourcesRef.current;
// Close connections for IDs no longer pending
for (const [id, source] of currentSources.entries()) {
if (!pendingIds.has(id)) {
source.close();
currentSources.delete(id);
}
}
// Open connections for newly pending IDs
for (const id of pendingIds) {
if (currentSources.has(id)) continue;
const url = apiClient.getGenerationStatusUrl(id);
const source = new EventSource(url);
source.onmessage = event => {
try {
const data: GenerationStatusEvent = JSON.parse(event.data);
if (data.status === 'completed') {
source.close();
currentSources.delete(id);
removePendingGeneration(id);
queryClient.refetchQueries({ queryKey: ['history'] });
const storyId = removePendingStoryAdd(id);
if (storyId) {
apiClient.addStoryItem(storyId, { generation_id: id })
.then(() => {
queryClient.invalidateQueries({ queryKey: ['stories'] });
toast({ title: 'Added to story', description: `Audio generated (${data.duration?.toFixed(2)}s)` });
});
}
if (autoplayRef.current && !isPlayingRef.current) {
const genAudioUrl = apiClient.getAudioUrl(id);
setAudioWithAutoPlay(genAudioUrl, id, '', '');
}
} else if (data.status === 'failed' || data.status === 'not_found') {
source.close();
currentSources.delete(id);
removePendingGeneration(id);
removePendingStoryAdd(id);
queryClient.refetchQueries({ queryKey: ['history'] });
toast({
title: data.status === 'not_found' ? 'Generation not found' : 'Generation failed',
description: data.error ?? 'An error occurred',
variant: 'destructive',
});
}
} catch {
// Ignore malformed messages
}
};
source.onerror = () => {
source.close();
currentSources.delete(id);
removePendingGeneration(id);
queryClient.refetchQueries({ queryKey: ['history'] });
};
currentSources.set(id, source);
}
}, [
pendingIds,
removePendingGeneration,
removePendingStoryAdd,
queryClient,
toast,
setAudioWithAutoPlay,
]);
}
The hook uses eventSourcesRef to prevent duplicate connections for the same generation ID. Connection cleanup occurs both when IDs leave the pending set and when the component unmounts, preventing memory leaks.
API Client Configuration
The ApiClient class centralizes URL construction for SSE endpoints and audio retrieval.
// app/src/lib/api/client.ts
class ApiClient {
getGenerationStatusUrl(generationId: string): string {
return `${this.getBaseUrl()}/generate/${generationId}/status`;
}
getAudioUrl(audioId: string): string {
return `${this.getBaseUrl()}/audio/${audioId}`;
}
// ... additional methods
}
This abstraction ensures consistent base URL handling across environments and simplifies endpoint management within React components.
Server-Side Implementation
Status Streaming Endpoint
The FastAPI route in backend/routes/generations.py implements the SSE protocol using Python's async generators and FastAPI's StreamingResponse.
# backend/routes/generations.py
@router.get("/generate/{generation_id}/status")
async def get_generation_status(generation_id: str, db: Session = Depends(get_db)):
"""SSE endpoint that streams generation status updates."""
import json
async def event_stream():
try:
while True:
db.expire_all()
gen = db.query(DBGeneration).filter_by(id=generation_id).first()
if not gen:
yield f"data: {json.dumps({'status': 'not_found', 'id': generation_id})}\n\n"
return
payload = {
"id": gen.id,
"status": gen.status or "completed",
"duration": gen.duration,
"error": gen.error,
}
yield f"data: {json.dumps(payload)}\n\n"
if (gen.status or "completed") in ("completed", "failed"):
return
await asyncio.sleep(1)
except (BrokenPipeError, ConnectionResetError, asyncio.CancelledError):
logger.debug("SSE client disconnected for generation %s", generation_id)
return StreamingResponse(
event_stream(),
media_type="text/event-stream",
headers={
"Cache-Control": "no-cache",
"Connection": "keep-alive",
"X-Accel-Buffering": "no",
},
)
The coroutine queries the database every second, yielding properly formatted SSE data lines (data: {...}\n\n). The X-Accel-Buffering: no header prevents reverse proxies from buffering the stream, ensuring immediate delivery to clients. Disconnection exceptions are caught silently to prevent server-side errors when users close browser tabs.
Direct Audio Streaming
For scenarios requiring immediate playback without disk persistence, Voicebox exposes a secondary streaming endpoint that transmits raw WAV bytes in chunks.
# backend/routes/generations.py
@router.post("/generate/stream")
async def stream_speech(
data: models.GenerationRequest,
db: Session = Depends(get_db),
):
"""Generate speech and stream the WAV audio directly without saving to disk."""
# ... model loading and audio generation logic ...
wav_bytes = tts.audio_to_wav_bytes(audio, sample_rate)
async def _wav_stream():
try:
chunk_size = 64 * 1024
for i in range(0, len(wav_bytes), chunk_size):
yield wav_bytes[i: i + chunk_size]
except (BrokenPipeError, ConnectionResetError, asyncio.CancelledError):
logger.debug("Client disconnected during audio stream")
return StreamingResponse(
_wav_stream(),
media_type="audio/wav",
headers={"Content-Disposition": 'attachment; filename="speech.wav"'},
)
The 64 KB chunk size balances memory efficiency with streaming performance, allowing the client to begin playback while the server is still transmitting data.
Complete Implementation Example
Integrating async generation into a React component requires coordinating the generation request with the progress hook.
import { useGenerationStore } from '@/stores/generationStore';
import { apiClient } from '@/lib/api/client';
import { useGenerationProgress } from '@/hooks/useGenerationProgress';
export function GenerateButton({ text }: { text: string }) {
const addPending = useGenerationStore(s => s.addPendingGeneration);
useGenerationProgress(); // Registers SSE listeners automatically
const handleGenerate = async () => {
const resp = await apiClient.postGeneration({
text,
profile_id: "default"
});
addPending(resp.id); // Triggers SSE connection
};
return <button onClick={handleGenerate}>Generate Speech</button>;
}
When the user clicks the button, the component posts to the generation endpoint and immediately adds the returned ID to the pending store. The useGenerationProgress hook detects this new ID, opens an SSE connection, and manages all subsequent state updates including toast notifications, cache invalidation, and automatic audio playback.
Summary
- Voicebox uses native EventSource APIs to maintain persistent connections for real-time generation progress, eliminating the need for client-side polling.
- The
useGenerationProgresshook maintains a Map of active connections, ensuring one SSE stream per generation ID with automatic cleanup on completion or failure. - FastAPI's StreamingResponse powers the backend, emitting JSON status updates every second via the
/generate/{id}/statusendpoint until reaching terminal states. - Two streaming modes exist: SSE for progress metadata and direct binary streaming via
/generate/streamfor immediate audio playback without disk writes. - Robust error handling catches disconnections at both network and application layers, preventing resource leaks and ensuring clean UI state transitions.
Frequently Asked Questions
How does Voicebox handle SSE connection failures?
The implementation includes multiple safeguards against connection instability. In useGenerationProgress.ts, the source.onerror handler closes the EventSource, removes the ID from pending generations, and refetches the history query to sync state. On the backend, the event_stream coroutine catches BrokenPipeError, ConnectionResetError, and asyncio.CancelledError to handle abrupt client disconnections without crashing the server.
What is the polling interval for generation status updates?
The backend explicitly throttles database queries using await asyncio.sleep(1) in the event loop. This one-second interval balances real-time responsiveness with database load, as implemented in backend/routes/generations.py. The client receives updates as Server-Sent Events rather than through active polling, reducing network overhead compared to traditional REST polling architectures.
Can Voicebox stream audio before generation completes?
The current implementation supports two distinct streaming patterns. The /generate/stream endpoint streams raw WAV bytes in 64 KB chunks immediately after synthesis completes, but requires the full audio buffer to be ready first. True progressive streaming during generation would require chunked TTS output and WebSocket connections, which are not implemented in the current architecture according to the source code.
How does the frontend know when to play audio automatically?
Auto-playback relies on the autoplayOnGenerate setting stored in the server configuration store (useServerStore). When the SSE stream reports status: "completed", the useGenerationProgress hook checks autoplayRef.current and isPlayingRef.current before calling setAudioWithAutoPlay with the audio URL constructed via apiClient.getAudioUrl(id). This ensures audio only plays automatically if the user enabled the setting and no other audio is currently playing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →