Creating a Voice Assistant with ElevenLabs STT and TTS Integration with Claude
Build a low-latency voice assistant by chaining ElevenLabs speech-to-text, streaming Claude responses, and ElevenLabs text-to-speech to achieve sub-1.5 second time-to-first-audio.
The anthropics/claude-cookbooks repository provides a complete reference implementation for building real-time conversational agents. The Low-Latency Voice Assistant notebook in third_party/ElevenLabs/low_latency_stt_claude_tts.ipynb demonstrates how to integrate ElevenLabs STT and TTS services with Claude's streaming API to minimize perceived latency in voice interactions.
Architecture Overview
The reference implementation follows a five-stage pipeline that processes audio input and generates spoken responses with minimal delay:
- Audio Capture: Record or load audio buffers for transcription.
- Speech-to-Text: Convert audio to text using ElevenLabs'
scribe_v1model. - LLM Processing: Stream the transcription to Claude via
anthropic_client.messages.stream. - Text-to-Speech: Synthesize responses using ElevenLabs'
eleven_turbo_v2_5model. - Latency Measurement: Instrument each stage to benchmark performance.
This architecture achieves first-token latency of approximately 1.48 seconds by streaming Claude's output sentence-by-sentence rather than waiting for the complete response. According to the source code in anthropics/claude-cookbooks, the implementation uses modular client objects that make it easy to swap providers without restructuring the pipeline.
Environment Setup and Authentication
The notebook requires API keys for both services, loaded securely via environment variables rather than hard-coded strings. Create a .env file in your project root with the following variables:
ELEVENLABS_API_KEY=your_elevenlabs_key_here
ANTHROPIC_API_KEY=your_anthropic_key_here
Initialize the clients using the official Python SDKs:
from dotenv import load_dotenv
import os, anthropic, elevenlabs
load_dotenv()
elevenlabs_client = elevenlabs.ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
base_url="https://api.elevenlabs.io"
)
anthropic_client = anthropic.Anthropic(
api_key=os.getenv("ANTHROPIC_API_KEY")
)
Voice Selection
The implementation programmatically selects the first available voice from your ElevenLabs account. In third_party/ElevenLabs/low_latency_stt_claude_tts.ipynb, the code searches available voices and extracts the voice_id for subsequent TTS calls:
voices = elevenlabs_client.voices.search().voices
selected_voice = voices[0] # e.g., Rachel
VOICE_ID = selected_voice.voice_id
print(f"Selected voice: {selected_voice.name} – {VOICE_ID}")
Converting Speech to Text
The STT stage uses ElevenLabs' speech_to_text.convert method to transcribe audio buffers. The implementation wraps this call with timing instrumentation to measure latency:
import io, time
audio_data.seek(0) # rewind BytesIO buffer
start = time.time()
transcription = elevenlabs_client.speech_to_text.convert(
file=audio_data,
model_id="scribe_v1"
)
print(f"Transcribed text: {transcription.text}")
print(f"Transcription time: {time.time() - start:.2f}s")
The transcription.text field contains the recognized user utterance, which feeds directly into Claude's context window.
Streaming Responses from Claude
To minimize time-to-first-token, the implementation uses Claude's streaming API rather than synchronous calls. The anthropic_client.messages.stream context manager yields tokens as they arrive from the claude-haiku-4-5 model:
start = time.time()
first_token_time = None
full_response = ""
with anthropic_client.messages.stream(
model="claude-haiku-4-5",
max_tokens=1000,
temperature=0,
messages=[{"role": "user", "content": transcription.text}],
) as stream:
for token in stream.text_stream:
full_response += token
print(token, end="", flush=True)
if first_token_time is None:
first_token_time = time.time()
print(f"\nStreaming time to first token: {first_token_time - start:.2f}s")
This approach prevents the perceptible delay typical of wait-for-completion patterns, providing immediate text feedback that can stream to TTS engines.
Synthesizing Speech with TTS
The basic TTS implementation streams the complete Claude response to ElevenLabs' text_to_speech.stream method, capturing the first audio chunk timestamp:
audio_buf = io.BytesIO()
first_chunk = None
start = time.time()
audio_gen = elevenlabs_client.text_to_speech.stream(
voice_id=VOICE_ID,
output_format="mp3_44100_128",
text=full_response,
model_id="eleven_turbo_v2_5",
)
for chunk in audio_gen:
if first_chunk is None:
first_chunk = time.time()
audio_buf.write(chunk)
print(f"Streaming TTS first chunk: {first_chunk - start:.2f}s")
While this method works, it waits for Claude to finish generating the entire response before beginning audio synthesis. For lower latency, the notebook implements a sentence-level streaming approach.
Sentence-by-Sentence Streaming Optimization
The lowest-latency implementation processes Claude's output incrementally, sending complete sentences to the TTS engine as soon as punctuation marks (., !, ?) appear. This allows audio playback to begin while Claude is still generating the remainder of the response:
import re
sentence_pat = re.compile(r"[.!?]+")
sentence_buf = ""
first_audio = None
audio_parts = []
start = time.time()
with anthropic_client.messages.stream(
model="claude-haiku-4-5",
max_tokens=1000,
temperature=0,
messages=[{"role": "user", "content": transcription.text}],
) as stream:
for token in stream.text_stream:
print(token, end="", flush=True)
sentence_buf += token
if sentence_pat.search(sentence_buf):
sentences = sentence_pat.split(sentence_buf)
for s in sentences[:-1]:
s = s.strip()
if not s:
continue
tts_gen = elevenlabs_client.text_to_speech.stream(
voice_id=VOICE_ID,
output_format="mp3_44100_128",
text=s,
model_id="eleven_turbo_v2_5",
)
audio_chunk = io.BytesIO()
for chunk in tts_gen:
if first_audio is None:
first_audio = time.time()
audio_chunk.write(chunk)
audio_parts.append(audio_chunk.getvalue())
sentence_buf = sentences[-1]
print(f"Time to first audio: {first_audio - start:.2f}s")
This pattern achieves time-to-first-audio of approximately 1.48 seconds, significantly faster than wait-for-completion approaches. The notebook notes that production implementations should consider ElevenLabs' WebSocket API to maintain prosody across sentence boundaries without buffering.
Latency Measurement and Benchmarking
The reference implementation instruments three critical latency metrics:
- STT latency: Time from audio submission to transcription receipt
- LLM time-to-first-token: Time from API call to first Claude response token
- TTS latency: Time from text input to first audio chunk generation
These measurements use standard time.time() calls at each stage boundary, enabling developers to identify bottlenecks in their specific deployment environments. As implemented in anthropics/claude-cookbooks, this instrumentation requires no external dependencies beyond Python's standard library.
Summary
- Initialize clients for both ElevenLabs and Anthropic using environment variables stored in
.env, not hard-coded credentials. - Stream Claude responses using
anthropic_client.messages.streamto minimize time-to-first-token rather than waiting for complete responses. - Process TTS sentence-by-sentence by buffering tokens until punctuation appears, then immediately calling
elevenlabs_client.text_to_speech.streamfor completed sentences. - Measure latency at each pipeline stage using
time.time()to benchmark and optimize performance. - Reference the complete notebook at
third_party/ElevenLabs/low_latency_stt_claude_tts.ipynbfor the full working implementation including audio playback examples.
Frequently Asked Questions
How do I reduce latency in a Claude voice assistant?
Process Claude's output using the streaming API (anthropic_client.messages.stream) and synthesize speech sentence-by-sentence rather than waiting for the complete response. This architecture, as shown in the anthropics/claude-cookbooks implementation, reduces time-to-first-audio to approximately 1.48 seconds.
What ElevenLabs models work best for STT and TTS integration?
Use the scribe_v1 model for speech-to-text transcription and eleven_turbo_v2_5 for text-to-speech synthesis. These models provide the speed and quality necessary for real-time conversational agents.
Can I use this architecture with other LLM providers?
Yes. The notebook uses modular client objects that abstract the API calls. Replacing the anthropic.Anthropic client with another provider's SDK requires only updating the initialization and message creation logic, while the ElevenLabs STT/TTS components remain unchanged.
How should I handle API keys securely in production?
Load credentials from environment variables using python-dotenv or your deployment platform's secret management system. The reference implementation uses load_dotenv() to populate ELEVENLABS_API_KEY and ANTHROPIC_API_KEY from a .env file, keeping secrets out of source control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →