# How to Build Real-Time Voice AI Agents: A Complete Implementation Guide

> Learn to build real-time voice AI agents using streaming TTS, dual LLM agents, and vector databases. Get a complete implementation guide for interactive voice applications.

- Repository: [Shubham Saboo/awesome-llm-apps](https://github.com/shubhamsaboo/awesome-llm-apps)
- Tags: how-to-guide
- Published: 2026-02-16

---

**Building real-time voice AI agents involves combining vector database retrieval with OpenAI's streaming TTS API, orchestrated through dual LLM agents for answer generation and speech optimization, all wrapped in an async Streamlit interface.**

This guide walks through the production-ready implementation found in the `Shubhamsaboo/awesome-llm-apps` repository, specifically the `voice_rag_openaisdk` project. The system demonstrates how to create a **Retrieval-Augmented Generation (RAG)** voice agent that ingests PDF documents, retrieves relevant context using **FastEmbed** and **Qdrant**, and streams natural-sounding audio responses using OpenAI's `gpt-4o-mini-tts` model.

## Architecture Overview

The real-time voice AI agent follows a modular pipeline architecture designed for low-latency audio streaming. The core components include:

- **Document Ingestion Layer**: PDFs are processed using `PyPDFLoader` and chunked with `RecursiveCharacterTextSplitter`
- **Embedding & Storage**: Text chunks are embedded using **FastEmbed** and stored in **Qdrant** with cosine similarity
- **Retrieval Engine**: Query embeddings fetch the top-k most relevant document chunks
- **Dual-Agent Orchestration**: A **processor agent** generates concise answers while a **TTS agent** optimizes text for voice synthesis
- **Audio Streaming**: Asynchronous streaming to the browser using `AsyncOpenAI` with real-time PCM playback and MP3 download options

## Setting Up the Environment and Session State

The application uses Streamlit's session state to persist API credentials and agent instances across interactions. The `init_session_state()` function in [`voice_ai_agents/voice_rag_openaisdk/rag_voice.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/voice_ai_agents/voice_rag_openaisdk/rag_voice.py) initializes all required variables:

```python
def init_session_state() -> None:
    defaults = {
        "initialized": False,
        "qdrant_url": "",
        "qdrant_api_key": "",
        "openai_api_key": "",
        "setup_complete": False,
        "client": None,
        "embedding_model": None,
        "processor_agent": None,
        "tts_agent": None,
        "selected_voice": "coral",
        "processed_documents": []
    }
    for k, v in defaults.items():
        if k not in st.session_state:
            st.session_state[k] = v

```

The sidebar configuration collects credentials and allows voice selection from OpenAI's TTS catalog:

```python
def setup_sidebar() -> None:
    with st.sidebar:
        st.title("🔑 Configuration")
        st.session_state.qdrant_url = st.text_input("Qdrant URL", type="password")
        st.session_state.qdrant_api_key = st.text_input("Qdrant API Key", type="password")
        st.session_state.openai_api_key = st.text_input("OpenAI API Key", type="password")
        voices = ["alloy","ash","ballad","coral","echo","fable","onyx","nova","sage","shimmer","verse"]
        st.session_state.selected_voice = st.selectbox("Select Voice", options=voices,
                                                      index=voices.index(st.session_state.selected_voice))

```

## Document Ingestion and Vector Storage

When users upload PDFs, the `process_pdf()` function handles extraction and metadata enrichment:

```python
def process_pdf(file) -> List:
    with tempfile.NamedTemporaryFile(delete=False, suffix='.pdf') as tmp_file:
        tmp_file.write(file.getvalue())
        loader = PyPDFLoader(tmp_file.name)
        docs = loader.load()

        for doc in docs:
            doc.metadata.update({
                "source_type": "pdf",
                "file_name": file.name,
                "timestamp": datetime.now().isoformat()
            })

        splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
        return splitter.split_documents(docs)

```

The `store_embeddings()` function persists chunks to Qdrant using FastEmbed vectors:

```python
def store_embeddings(client, embedding_model, documents, collection_name):
    for doc in documents:
        emb = list(embedding_model.embed([doc.page_content]))[0]
        client.upsert(
            collection_name=collection_name,
            points=[
                models.PointStruct(
                    id=str(uuid.uuid4()),
                    vector=emb.tolist(),
                    payload={"content": doc.page_content, **doc.metadata}
                )
            ]
        )

```

## Initializing the Vector Database

The `setup_qdrant()` function establishes the connection and ensures the collection exists with proper vector dimensions:

```python
def setup_qdrant() -> Tuple[QdrantClient, TextEmbedding]:
    if not all([st.session_state.qdrant_url, st.session_state.qdrant_api_key]):
        raise ValueError("Qdrant credentials not provided")
    client = QdrantClient(url=st.session_state.qdrant_url,
                         api_key=st.session_state.qdrant_api_key)

    embedding_model = TextEmbedding()
    dim = len(list(embedding_model.embed(["test"]))[0])

    # Create collection if it does not exist

    try:
        client.create_collection(
            collection_name=COLLECTION_NAME,
            vectors_config=VectorParams(size=dim, distance=Distance.COSINE)
        )
    except Exception as e:
        if "already exists" not in str(e):
            raise e

    return client, embedding_model

```

## Dual-Agent Orchestration for Voice Optimization

The system uses two specialized agents defined in `setup_agents()`. The **processor agent** generates concise, factual answers optimized for spoken delivery, while the **TTS agent** refines the text for natural speech patterns:

```python
def setup_agents(openai_api_key: str) -> Tuple[Agent, Agent]:
    os.environ["OPENAI_API_KEY"] = openai_api_key

    processor_agent = Agent(
        name="Documentation Processor",
        instructions="""You are a helpful documentation assistant. ...
        Format your response in a way that's easy to speak out loud""",
        model="gpt-4o"
    )
    tts_agent = Agent(
        name="Text-to-Speech Agent",
        instructions="""You are a text-to-speech agent. ...
        Ensure the speech is clear and well‑articulated""",
        model="gpt-4o"
    )
    return processor_agent, tts_agent

```

## Real-Time Audio Streaming Pipeline

The `process_query()` function in [`rag_voice.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_voice.py) implements the complete async pipeline. It retrieves context from Qdrant, runs both agents, and streams audio using OpenAI's TTS endpoint:

```python
async def process_query(query, client, embedding_model, collection_name,
                        openai_api_key, voice):
    # 1️⃣ Embed & search

    query_emb = list(embedding_model.embed([query]))[0]
    search = client.query_points(collection_name=collection_name,
                                query=query_emb.tolist(),
                                limit=3, with_payload=True).points

    # 2️⃣ Build context

    context = "Based on the following documentation:\n\n"
    for r in search:
        payload = r.payload
        if payload:
            content = payload.get('content', '')
            source = payload.get('file_name', 'Unknown')
            context += f"From {source}:\n{content}\n\n"
    context += f"\nUser Question: {query}\n\nPlease provide a clear, concise answer that can be easily spoken out loud."

    # 3️⃣ Initialise agents if needed

    if not st.session_state.processor_agent or not st.session_state.tts_agent:
        proc, tts = setup_agents(openai_api_key)
        st.session_state.processor_agent, st.session_state.tts_agent = proc, tts

    # 4️⃣ Generate text answer

    processor_result = await Runner.run(st.session_state.processor_agent, context)
    text_response = processor_result.final_output

    # 5️⃣ Generate TTS instructions

    tts_result = await Runner.run(st.session_state.tts_agent, text_response)
    voice_instructions = tts_result.final_output

    # 6️⃣ Stream / save audio

    async_openai = AsyncOpenAI(api_key=openai_api_key)
    async with async_openai.audio.speech.with_streaming_response.create(
            model="gpt-4o-mini-tts",
            voice=voice,
            input=text_response,
            instructions=voice_instructions,
            response_format="pcm",
    ) as stream_response:
        await LocalAudioPlayer().play(stream_response)  # real‑time playback

        # MP3 version for download

        audio_resp = await async_openai.audio.speech.create(
            model="gpt-4o-mini-tts",
            voice=voice,
            input=text_response,
            instructions=voice_instructions,
            response_format="mp3"
        )
        audio_path = os.path.join(tempfile.gettempdir(),
                                  f"response_{uuid.uuid4()}.mp3")
        with open(audio_path, "wb") as f:
            f.write(audio_resp.content)

    return {
        "status": "success",
        "text_response": text_response,
        "audio_path": audio_path,
        "sources": [p.payload.get('file_name') for p in search if p.payload]
    }

```

## Building the Streamlit Interface

The `main()` function in [`rag_voice.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_voice.py) orchestrates the complete user flow, from file upload to audio playback:

```python
def main() -> None:
    st.set_page_config(page_title="Voice RAG Agent", layout="wide")
    init_session_state()
    setup_sidebar()

    st.title("🎙️ Voice RAG Agent")
    uploaded_file = st.file_uploader("Upload PDF", type=["pdf"])

    if uploaded_file:
        # (process PDF, store embeddings – see section 2)

    query = st.text_input("What would you like to know ...",
                          disabled=not st.session_state.setup_complete)

    if query and st.session_state.setup_complete:
        result = asyncio.run(process_query(
            query,
            st.session_state.client,
            st.session_state.embedding_model,
            COLLECTION_NAME,
            st.session_state.openai_api_key,
            st.session_state.selected_voice
        ))
        # Render text, audio player, download button, sources...

```

## Alternative Voice Agent Patterns

The repository contains additional implementations that demonstrate different approaches to building real-time voice AI agents.

### Customer Support Voice Agent

Located in [`voice_ai_agents/customer_support_voice_agent/customer_support_voice_agent.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/voice_ai_agents/customer_support_voice_agent/customer_support_voice_agent.py), this variant uses **Firecrawl** to crawl live documentation websites rather than processing static PDFs. It maintains the same dual-agent architecture (processor + TTS) but dynamically builds the knowledge base from web content, making it ideal for support teams needing voice-enabled access to constantly evolving documentation.

### AI Audio Tour Agent

Found in [`voice_ai_agents/ai_audio_tour_agent/ai_audio_tour_agent.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/voice_ai_agents/ai_audio_tour_agent/ai_audio_tour_agent.py), this implementation demonstrates single-shot audio generation without retrieval. It generates location-based tour content using GPT-4o and immediately converts it to speech using `AsyncOpenAI.audio.speech.create()`, producing downloadable MP3 files. This pattern works well for content creation workflows rather than interactive RAG applications.

## Running the Voice RAG Agent Locally

To deploy this real-time voice AI agent on your machine:

1. **Clone the repository**:
   ```bash
   git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git
   cd awesome-llm-apps/voice_ai_agents/voice_rag_openaisdk
   ```

2. **Install dependencies**:
   ```bash
   pip install -r requirements.txt
   ```

3. **Configure API keys** via environment variables or the Streamlit sidebar:
   ```text
   OPENAI_API_KEY=sk-xxxx
   QDRANT_URL=https://xxxx.qdrant.io
   QDRANT_API_KEY=xxxxxxxx
   ```

4. **Launch the application**:
   ```bash
   streamlit run rag_voice.py
   ```

5. **Interact** by uploading PDF documents, entering questions, and receiving instant spoken responses with downloadable audio files.

## Summary

Building production-ready real-time voice AI agents requires orchestrating several key components:

- **Vector retrieval infrastructure** using FastEmbed for document embeddings and Qdrant for semantic search enables context-aware responses
- **Dual-agent architecture** separates content generation (processor agent) from speech optimization (TTS agent), producing natural-sounding voice output
- **Asynchronous streaming** via `AsyncOpenAI` and the `gpt-4o-mini-tts` model delivers sub-second audio playback while maintaining MP3 download capabilities
- **Streamlit integration** provides the session management and UI components necessary for interactive document upload and real-time interaction

## Frequently Asked Questions

### What makes this voice AI agent "real-time" rather than just text-to-speech?

The implementation achieves real-time performance through **asynchronous audio streaming** using `AsyncOpenAI.audio.speech.with_streaming_response.create()` with PCM response format, which streams audio bytes to `LocalAudioPlayer()` for immediate playback while the request is still processing. This contrasts with traditional TTS approaches that wait for the entire MP3 file to generate before playing.

### Can I replace Qdrant with Pinecone or Weaviate for the vector storage?

Yes, the embedding and storage logic in `store_embeddings()` and `setup_qdrant()` uses standard vector operations that translate to other providers. You would need to modify the client initialization in `setup_qdrant()` to use your preferred vector database's Python client, adjust the collection creation parameters, and update the query logic in `process_query()` to use the specific search syntax of your chosen database.

### How do I customize the voice personality beyond the standard OpenAI voices?

While the voice selection is limited to OpenAI's catalog (alloy, coral, echo, etc.), you can customize personality through the **TTS agent instructions** in `setup_agents()`. By modifying the `tts_agent` system prompt, you can instruct the model to format output with specific emotional tones, pacing cues, or speaking styles (e.g., "speak enthusiastically" or "use pauses for emphasis") which the `gpt-4o-mini-tts` model interprets during speech synthesis.

### What are the API costs associated with running this voice agent?

The implementation uses three billable OpenAI services: **GPT-4o** for the processor and TTS agents (text generation), **GPT-4o-mini-tts** for speech synthesis (audio generation), and **text-embedding-3-small** (via FastEmbed) for document embeddings. Qdrant Cloud offers a free tier sufficient for development. Actual costs depend on document volume and query frequency, but the use of `gpt-4o-mini-tts` specifically optimizes for lower latency and cost compared to standard TTS models while maintaining quality.