How the Gods Eye View Voice System Uses the OpenAI Realtime API: A WebRTC Implementation Guide

The voice subsystem in bilawalsidhu/gods-eye-view implements a WebRTC-based client that acquires ephemeral tokens, establishes a peer-to-peer connection with OpenAI's Realtime API, and manages bidirectional audio streaming with built-in cost capping.

The repository implements a complete client-side voice AI loop that streams microphone audio to OpenAI and receives spoken responses alongside structured function calls. This architecture leverages the OpenAI Realtime API's native WebRTC support to minimize latency while maintaining strict cost controls directly in the browser.

Token Acquisition and Authentication

Before establishing any audio connection, the system must secure a short-lived access token. The development server exposes a dedicated endpoint at /api/realtime/token defined in vite.config.js:

// vite.config.js - Token endpoint handler
server: {
  '/api/realtime/token': async (req, res) => {
    // Proxies to OpenAI with model tier configuration
    // Returns { token: string, model: string }
  }
}

When GevRealtimeController.start() initializes a session, it fetches this token via an internal HTTP request. The token embeds the requested model tier and expires quickly, ensuring credentials never hardcode into the client bundle. This token then authenticates all subsequent WebRTC signaling requests to https://api.openai.com/v1/realtime/calls.

WebRTC Connection Setup

The core controller in src/voice/gevRealtime.js orchestrates the peer connection lifecycle. The start() method creates a standard RTCPeerConnection, adds the user's microphone tracks, and opens a data channel named oai-events for JSON signaling:

// src/voice/gevRealtime.js - Connection initialization
handleStart() {
  this.peerConnection = new RTCPeerConnection({
    iceServers: [{ urls: 'stun:stun.l.google.com:19302' }]
  });
  
  // Add microphone tracks
  this.localStream.getTracks().forEach(track => {
    this.peerConnection.addTrack(track, this.localStream);
  });
  
  // Establish data channel for events
  this.dataChannel = this.peerConnection.createDataChannel('oai-events');
}

The controller then generates an SDP offer and POSTs it to the OpenAI Realtime endpoint with the bearer token in the Authorization header. Upon receiving the SDP answer, the browser sets it as the remote description, completing the bidirectional media and data path.

Audio Routing and Visualization

Once connected, audio flows through two distinct paths:

  • Upstream: The microphone stream captured via getUserMedia() transmits directly through the RTCPeerConnection to OpenAI's servers
  • Downstream: An <audio> element receives the remote stream and plays the assistant's spoken responses

The controller initializes visual feedback through the Web Audio API. Functions like startVoiceVisualizer() and startAssistantVoiceVisualizer() analyze audio levels to render real-time waveforms, providing visual confirmation that the system is listening or speaking.

Realtime Event Handling on the Data Channel

All communication beyond raw audio occurs as JSON-encoded events over the oai-events data channel. The handleRealtimeEvent method in gevRealtime.js parses these messages and dispatches them to specific handlers:

  • input_audio_buffer.speech_started: Marks the beginning of a user turn, triggering UI state changes to indicate active listening
  • response.done: Contains usage metadata including token counts; the controller forwards this payload to the cost tracker for billing calculations
  • Function call events: Extracted events like function_call are deduplicated and routed to the action runner in src/voice/gevActions.js

When OpenAI requests a tool execution, the controller runs the appropriate function and returns results via a function_call_output event:

// Example: Handling radio control from OpenAI
async handleFunctionCall(event) {
  const { name, arguments: args } = event;
  
  // Execute via action runner
  const result = await this.actionRunner.execute(name, args);
  
  // Return result to OpenAI
  this.sendRealtimeEvent({
    type: 'conversation.item.create',
    item: {
      type: 'function_call_output',
      call_id: event.call_id,
      output: JSON.stringify(result)
    }
  });
}

Cost Tracking and Safety Limits

To prevent runaway spending, src/voice/voiceCost.js implements a VoiceCostTracker class that monitors every response.done event. The tracker maintains a registry of model-specific pricing rates and converts token usage to USD in real-time:

// src/voice/voiceCost.js - Cost tracking implementation
class VoiceCostTracker {
  updateFromResponse(usage) {
    const cost = (usage.input_tokens * INPUT_RATE) + 
                 (usage.output_tokens * OUTPUT_RATE);
    this.totalSpend += cost;
    return this.checkThresholds();
  }
  
  checkThresholds() {
    if (this.totalSpend >= HARD_CAP) return { level: 'cap', action: 'stop' };
    if (this.totalSpend >= SOFT_WARN) return { level: 'warn' };
    return { level: 'ok' };
  }
}

The system enforces two configurable thresholds: a soft warning and a hard cap. When the cap triggers, the controller automatically invokes stop() to terminate the session, ensuring voice interactions never exceed budget constraints.

Session Lifecycle and Graceful Teardown

Terminating a session requires careful cleanup to release hardware resources and network connections. The stop() method in gevRealtimeController performs the following sequence:

  1. Aborts any pending function calls currently executing in the action runner
  2. Closes the oai-events data channel to stop event processing
  3. Tears down the RTCPeerConnection and releases ICE gathering resources
  4. Stops all microphone tracks to turn off the recording indicator
  5. Resets UI state to idle, clearing visualizers and status messages

This systematic teardown prevents memory leaks and ensures the microphone permission indicator disappears immediately when the user ends the session.

Code Examples

Starting a Voice Session

Initialize the controller and begin streaming with push-to-talk disabled for hands-free operation:

import { initGevVoiceCommands } from './voice/gevRealtime.js';

const controller = initGevVoiceCommands({ 
  viewer, 
  styleManager, 
  dataManager 
});

// Begin session: fetches token, establishes WebRTC, starts microphone
controller.start({ pushToTalk: false });

Sending Text Commands

While the voice session remains active, inject text-based commands directly into the conversation:

controller.sendTextCommand('Show me the traffic in Osaka at 5pm');

This method creates a conversation.item.create event on the data channel and queues a response generation from the Realtime API.

Handling Function Calls

The action runner implements domain-specific tools that OpenAI can invoke. For example, radio control:

// src/voice/gevActions.js
export async function control_radio({ action }) {
  if (action === 'play') {
    await radioLayer.playback();
    return { 
      ok: true, 
      action, 
      radioPlaybackRequested: true 
    };
  }
  // Handle pause, stop, volume adjustments...
}

Monitoring Session Costs

Access real-time spending data through the cost tracker:

const snapshot = controller.costTracker.state();
console.log(`Current spend: ${snapshot.display}`); // "~$0.42"

if (snapshot.level === 'cap') {
  console.warn('Spend cap reached – session stopping');
  controller.stop();
}

Summary

  • Token-based authentication: The system proxys token requests through /api/realtime/token in vite.config.js to avoid exposing API keys in the client
  • Native WebRTC integration: GevRealtimeController manages RTCPeerConnection, SDP exchange with https://api.openai.com/v1/realtime/calls, and the oai-events data channel
  • Bidirectional audio: Microphone streams upload via WebRTC while assistant responses play through a standard HTML audio element
  • Event-driven architecture: The controller handles input_audio_buffer.speech_started, response.done, and function_call events to coordinate the conversation flow
  • Cost safety: VoiceCostTracker in voiceCost.js monitors usage from each response and enforces hard spending caps to prevent budget overruns
  • Clean teardown: The stop() method systematically releases WebRTC resources, data channels, and microphone tracks

Frequently Asked Questions

How does the voice system authenticate with the OpenAI Realtime API?

The system uses ephemeral bearer tokens obtained through a backend proxy endpoint. When a session starts, the client requests a token from /api/realtime/token, which the development server defined in vite.config.js generates by forwarding the request to OpenAI with the appropriate model tier. These short-lived tokens authenticate the WebRTC handshake without exposing permanent API credentials in browser code.

What WebRTC components handle the audio streaming?

The implementation uses a standard RTCPeerConnection configured with STUN servers for NAT traversal. The controller adds local microphone tracks to the connection for upstream audio, while the remote audio stream attaches to an HTML <audio> element for playback. A dedicated data channel named oai-events carries JSON signaling events alongside the media stream.

How are function calls from the Realtime API executed?

When the controller receives a function_call event over the data channel, it extracts the function name and arguments, then dispatches them to the action runner implemented in src/voice/gevActions.js. After executing the requested operation (such as controlling radio playback or toggling layer visibility), the controller formats the result as a function_call_output event and transmits it back to OpenAI to continue the conversation.

How does the system prevent excessive billing during voice sessions?

A VoiceCostTracker instance monitors token usage from every response.done event and converts it to USD using model-specific rates. The tracker maintains soft warning and hard cap thresholds. When spending reaches the hard cap, the controller automatically triggers stop() to terminate the WebRTC connection and end the session immediately, ensuring costs remain bounded regardless of conversation length.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →