How to Build Real-Time Voice AI Agents: A Complete Implementation Guide
Building real-time voice AI agents involves combining vector database retrieval with OpenAI's streaming TTS API, orchestrated through dual LLM agents for answer generation and speech optimization, all wrapped in an async Streamlit interface.
This guide walks through the production-ready implementation found in the Shubhamsaboo/awesome-llm-apps repository, specifically the voice_rag_openaisdk project. The system demonstrates how to create a Retrieval-Augmented Generation (RAG) voice agent that ingests PDF documents, retrieves relevant context using FastEmbed and Qdrant, and streams natural-sounding audio responses using OpenAI's gpt-4o-mini-tts model.
Architecture Overview
The real-time voice AI agent follows a modular pipeline architecture designed for low-latency audio streaming. The core components include:
- Document Ingestion Layer: PDFs are processed using
PyPDFLoaderand chunked withRecursiveCharacterTextSplitter - Embedding & Storage: Text chunks are embedded using FastEmbed and stored in Qdrant with cosine similarity
- Retrieval Engine: Query embeddings fetch the top-k most relevant document chunks
- Dual-Agent Orchestration: A processor agent generates concise answers while a TTS agent optimizes text for voice synthesis
- Audio Streaming: Asynchronous streaming to the browser using
AsyncOpenAIwith real-time PCM playback and MP3 download options
Setting Up the Environment and Session State
The application uses Streamlit's session state to persist API credentials and agent instances across interactions. The init_session_state() function in voice_ai_agents/voice_rag_openaisdk/rag_voice.py initializes all required variables:
def init_session_state() -> None:
defaults = {
"initialized": False,
"qdrant_url": "",
"qdrant_api_key": "",
"openai_api_key": "",
"setup_complete": False,
"client": None,
"embedding_model": None,
"processor_agent": None,
"tts_agent": None,
"selected_voice": "coral",
"processed_documents": []
}
for k, v in defaults.items():
if k not in st.session_state:
st.session_state[k] = v
The sidebar configuration collects credentials and allows voice selection from OpenAI's TTS catalog:
def setup_sidebar() -> None:
with st.sidebar:
st.title("🔑 Configuration")
st.session_state.qdrant_url = st.text_input("Qdrant URL", type="password")
st.session_state.qdrant_api_key = st.text_input("Qdrant API Key", type="password")
st.session_state.openai_api_key = st.text_input("OpenAI API Key", type="password")
voices = ["alloy","ash","ballad","coral","echo","fable","onyx","nova","sage","shimmer","verse"]
st.session_state.selected_voice = st.selectbox("Select Voice", options=voices,
index=voices.index(st.session_state.selected_voice))
Document Ingestion and Vector Storage
When users upload PDFs, the process_pdf() function handles extraction and metadata enrichment:
def process_pdf(file) -> List:
with tempfile.NamedTemporaryFile(delete=False, suffix='.pdf') as tmp_file:
tmp_file.write(file.getvalue())
loader = PyPDFLoader(tmp_file.name)
docs = loader.load()
for doc in docs:
doc.metadata.update({
"source_type": "pdf",
"file_name": file.name,
"timestamp": datetime.now().isoformat()
})
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
return splitter.split_documents(docs)
The store_embeddings() function persists chunks to Qdrant using FastEmbed vectors:
def store_embeddings(client, embedding_model, documents, collection_name):
for doc in documents:
emb = list(embedding_model.embed([doc.page_content]))[0]
client.upsert(
collection_name=collection_name,
points=[
models.PointStruct(
id=str(uuid.uuid4()),
vector=emb.tolist(),
payload={"content": doc.page_content, **doc.metadata}
)
]
)
Initializing the Vector Database
The setup_qdrant() function establishes the connection and ensures the collection exists with proper vector dimensions:
def setup_qdrant() -> Tuple[QdrantClient, TextEmbedding]:
if not all([st.session_state.qdrant_url, st.session_state.qdrant_api_key]):
raise ValueError("Qdrant credentials not provided")
client = QdrantClient(url=st.session_state.qdrant_url,
api_key=st.session_state.qdrant_api_key)
embedding_model = TextEmbedding()
dim = len(list(embedding_model.embed(["test"]))[0])
# Create collection if it does not exist
try:
client.create_collection(
collection_name=COLLECTION_NAME,
vectors_config=VectorParams(size=dim, distance=Distance.COSINE)
)
except Exception as e:
if "already exists" not in str(e):
raise e
return client, embedding_model
Dual-Agent Orchestration for Voice Optimization
The system uses two specialized agents defined in setup_agents(). The processor agent generates concise, factual answers optimized for spoken delivery, while the TTS agent refines the text for natural speech patterns:
def setup_agents(openai_api_key: str) -> Tuple[Agent, Agent]:
os.environ["OPENAI_API_KEY"] = openai_api_key
processor_agent = Agent(
name="Documentation Processor",
instructions="""You are a helpful documentation assistant. ...
Format your response in a way that's easy to speak out loud""",
model="gpt-4o"
)
tts_agent = Agent(
name="Text-to-Speech Agent",
instructions="""You are a text-to-speech agent. ...
Ensure the speech is clear and well‑articulated""",
model="gpt-4o"
)
return processor_agent, tts_agent
Real-Time Audio Streaming Pipeline
The process_query() function in rag_voice.py implements the complete async pipeline. It retrieves context from Qdrant, runs both agents, and streams audio using OpenAI's TTS endpoint:
async def process_query(query, client, embedding_model, collection_name,
openai_api_key, voice):
# 1️⃣ Embed & search
query_emb = list(embedding_model.embed([query]))[0]
search = client.query_points(collection_name=collection_name,
query=query_emb.tolist(),
limit=3, with_payload=True).points
# 2️⃣ Build context
context = "Based on the following documentation:\n\n"
for r in search:
payload = r.payload
if payload:
content = payload.get('content', '')
source = payload.get('file_name', 'Unknown')
context += f"From {source}:\n{content}\n\n"
context += f"\nUser Question: {query}\n\nPlease provide a clear, concise answer that can be easily spoken out loud."
# 3️⃣ Initialise agents if needed
if not st.session_state.processor_agent or not st.session_state.tts_agent:
proc, tts = setup_agents(openai_api_key)
st.session_state.processor_agent, st.session_state.tts_agent = proc, tts
# 4️⃣ Generate text answer
processor_result = await Runner.run(st.session_state.processor_agent, context)
text_response = processor_result.final_output
# 5️⃣ Generate TTS instructions
tts_result = await Runner.run(st.session_state.tts_agent, text_response)
voice_instructions = tts_result.final_output
# 6️⃣ Stream / save audio
async_openai = AsyncOpenAI(api_key=openai_api_key)
async with async_openai.audio.speech.with_streaming_response.create(
model="gpt-4o-mini-tts",
voice=voice,
input=text_response,
instructions=voice_instructions,
response_format="pcm",
) as stream_response:
await LocalAudioPlayer().play(stream_response) # real‑time playback
# MP3 version for download
audio_resp = await async_openai.audio.speech.create(
model="gpt-4o-mini-tts",
voice=voice,
input=text_response,
instructions=voice_instructions,
response_format="mp3"
)
audio_path = os.path.join(tempfile.gettempdir(),
f"response_{uuid.uuid4()}.mp3")
with open(audio_path, "wb") as f:
f.write(audio_resp.content)
return {
"status": "success",
"text_response": text_response,
"audio_path": audio_path,
"sources": [p.payload.get('file_name') for p in search if p.payload]
}
Building the Streamlit Interface
The main() function in rag_voice.py orchestrates the complete user flow, from file upload to audio playback:
def main() -> None:
st.set_page_config(page_title="Voice RAG Agent", layout="wide")
init_session_state()
setup_sidebar()
st.title("🎙️ Voice RAG Agent")
uploaded_file = st.file_uploader("Upload PDF", type=["pdf"])
if uploaded_file:
# (process PDF, store embeddings – see section 2)
query = st.text_input("What would you like to know ...",
disabled=not st.session_state.setup_complete)
if query and st.session_state.setup_complete:
result = asyncio.run(process_query(
query,
st.session_state.client,
st.session_state.embedding_model,
COLLECTION_NAME,
st.session_state.openai_api_key,
st.session_state.selected_voice
))
# Render text, audio player, download button, sources...
Alternative Voice Agent Patterns
The repository contains additional implementations that demonstrate different approaches to building real-time voice AI agents.
Customer Support Voice Agent
Located in voice_ai_agents/customer_support_voice_agent/customer_support_voice_agent.py, this variant uses Firecrawl to crawl live documentation websites rather than processing static PDFs. It maintains the same dual-agent architecture (processor + TTS) but dynamically builds the knowledge base from web content, making it ideal for support teams needing voice-enabled access to constantly evolving documentation.
AI Audio Tour Agent
Found in voice_ai_agents/ai_audio_tour_agent/ai_audio_tour_agent.py, this implementation demonstrates single-shot audio generation without retrieval. It generates location-based tour content using GPT-4o and immediately converts it to speech using AsyncOpenAI.audio.speech.create(), producing downloadable MP3 files. This pattern works well for content creation workflows rather than interactive RAG applications.
Running the Voice RAG Agent Locally
To deploy this real-time voice AI agent on your machine:
-
Clone the repository:
git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git cd awesome-llm-apps/voice_ai_agents/voice_rag_openaisdk -
Install dependencies:
pip install -r requirements.txt -
Configure API keys via environment variables or the Streamlit sidebar:
OPENAI_API_KEY=sk-xxxx QDRANT_URL=https://xxxx.qdrant.io QDRANT_API_KEY=xxxxxxxx -
Launch the application:
streamlit run rag_voice.py -
Interact by uploading PDF documents, entering questions, and receiving instant spoken responses with downloadable audio files.
Summary
Building production-ready real-time voice AI agents requires orchestrating several key components:
- Vector retrieval infrastructure using FastEmbed for document embeddings and Qdrant for semantic search enables context-aware responses
- Dual-agent architecture separates content generation (processor agent) from speech optimization (TTS agent), producing natural-sounding voice output
- Asynchronous streaming via
AsyncOpenAIand thegpt-4o-mini-ttsmodel delivers sub-second audio playback while maintaining MP3 download capabilities - Streamlit integration provides the session management and UI components necessary for interactive document upload and real-time interaction
Frequently Asked Questions
What makes this voice AI agent "real-time" rather than just text-to-speech?
The implementation achieves real-time performance through asynchronous audio streaming using AsyncOpenAI.audio.speech.with_streaming_response.create() with PCM response format, which streams audio bytes to LocalAudioPlayer() for immediate playback while the request is still processing. This contrasts with traditional TTS approaches that wait for the entire MP3 file to generate before playing.
Can I replace Qdrant with Pinecone or Weaviate for the vector storage?
Yes, the embedding and storage logic in store_embeddings() and setup_qdrant() uses standard vector operations that translate to other providers. You would need to modify the client initialization in setup_qdrant() to use your preferred vector database's Python client, adjust the collection creation parameters, and update the query logic in process_query() to use the specific search syntax of your chosen database.
How do I customize the voice personality beyond the standard OpenAI voices?
While the voice selection is limited to OpenAI's catalog (alloy, coral, echo, etc.), you can customize personality through the TTS agent instructions in setup_agents(). By modifying the tts_agent system prompt, you can instruct the model to format output with specific emotional tones, pacing cues, or speaking styles (e.g., "speak enthusiastically" or "use pauses for emphasis") which the gpt-4o-mini-tts model interprets during speech synthesis.
What are the API costs associated with running this voice agent?
The implementation uses three billable OpenAI services: GPT-4o for the processor and TTS agents (text generation), GPT-4o-mini-tts for speech synthesis (audio generation), and text-embedding-3-small (via FastEmbed) for document embeddings. Qdrant Cloud offers a free tier sufficient for development. Actual costs depend on document volume and query frequency, but the use of gpt-4o-mini-tts specifically optimizes for lower latency and cost compared to standard TTS models while maintaining quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →