# How the Gods Eye View Voice System Uses the OpenAI Realtime API: A WebRTC Implementation Guide

> Explore the Gods Eye View voice system's WebRTC implementation of the OpenAI Realtime API for bidirectional audio streaming with cost capping. Learn how it works.

- Repository: [Bilawal Sidhu/gods-eye-view](https://github.com/bilawalsidhu/gods-eye-view)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The voice subsystem in `bilawalsidhu/gods-eye-view` implements a WebRTC-based client that acquires ephemeral tokens, establishes a peer-to-peer connection with OpenAI's Realtime API, and manages bidirectional audio streaming with built-in cost capping.**

The repository implements a complete client-side voice AI loop that streams microphone audio to OpenAI and receives spoken responses alongside structured function calls. This architecture leverages the OpenAI Realtime API's native WebRTC support to minimize latency while maintaining strict cost controls directly in the browser.

## Token Acquisition and Authentication

Before establishing any audio connection, the system must secure a short-lived access token. The development server exposes a dedicated endpoint at `/api/realtime/token` defined in [`vite.config.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/vite.config.js):

```javascript
// vite.config.js - Token endpoint handler
server: {
  '/api/realtime/token': async (req, res) => {
    // Proxies to OpenAI with model tier configuration
    // Returns { token: string, model: string }
  }
}

```

When `GevRealtimeController.start()` initializes a session, it fetches this token via an internal HTTP request. The token embeds the requested model tier and expires quickly, ensuring credentials never hardcode into the client bundle. This token then authenticates all subsequent WebRTC signaling requests to `https://api.openai.com/v1/realtime/calls`.

## WebRTC Connection Setup

The core controller in [`src/voice/gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevRealtime.js) orchestrates the peer connection lifecycle. The `start()` method creates a standard `RTCPeerConnection`, adds the user's microphone tracks, and opens a data channel named `oai-events` for JSON signaling:

```javascript
// src/voice/gevRealtime.js - Connection initialization
handleStart() {
  this.peerConnection = new RTCPeerConnection({
    iceServers: [{ urls: 'stun:stun.l.google.com:19302' }]
  });
  
  // Add microphone tracks
  this.localStream.getTracks().forEach(track => {
    this.peerConnection.addTrack(track, this.localStream);
  });
  
  // Establish data channel for events
  this.dataChannel = this.peerConnection.createDataChannel('oai-events');
}

```

The controller then generates an SDP offer and POSTs it to the OpenAI Realtime endpoint with the bearer token in the `Authorization` header. Upon receiving the SDP answer, the browser sets it as the remote description, completing the bidirectional media and data path.

## Audio Routing and Visualization

Once connected, audio flows through two distinct paths:

- **Upstream**: The microphone stream captured via `getUserMedia()` transmits directly through the `RTCPeerConnection` to OpenAI's servers
- **Downstream**: An `<audio>` element receives the remote stream and plays the assistant's spoken responses

The controller initializes visual feedback through the Web Audio API. Functions like `startVoiceVisualizer()` and `startAssistantVoiceVisualizer()` analyze audio levels to render real-time waveforms, providing visual confirmation that the system is listening or speaking.

## Realtime Event Handling on the Data Channel

All communication beyond raw audio occurs as JSON-encoded events over the `oai-events` data channel. The `handleRealtimeEvent` method in [`gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/gevRealtime.js) parses these messages and dispatches them to specific handlers:

- **`input_audio_buffer.speech_started`**: Marks the beginning of a user turn, triggering UI state changes to indicate active listening
- **`response.done`**: Contains usage metadata including token counts; the controller forwards this payload to the cost tracker for billing calculations
- **Function call events**: Extracted events like `function_call` are deduplicated and routed to the action runner in [`src/voice/gevActions.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevActions.js)

When OpenAI requests a tool execution, the controller runs the appropriate function and returns results via a `function_call_output` event:

```javascript
// Example: Handling radio control from OpenAI
async handleFunctionCall(event) {
  const { name, arguments: args } = event;
  
  // Execute via action runner
  const result = await this.actionRunner.execute(name, args);
  
  // Return result to OpenAI
  this.sendRealtimeEvent({
    type: 'conversation.item.create',
    item: {
      type: 'function_call_output',
      call_id: event.call_id,
      output: JSON.stringify(result)
    }
  });
}

```

## Cost Tracking and Safety Limits

To prevent runaway spending, [`src/voice/voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/voiceCost.js) implements a `VoiceCostTracker` class that monitors every `response.done` event. The tracker maintains a registry of model-specific pricing rates and converts token usage to USD in real-time:

```javascript
// src/voice/voiceCost.js - Cost tracking implementation
class VoiceCostTracker {
  updateFromResponse(usage) {
    const cost = (usage.input_tokens * INPUT_RATE) + 
                 (usage.output_tokens * OUTPUT_RATE);
    this.totalSpend += cost;
    return this.checkThresholds();
  }
  
  checkThresholds() {
    if (this.totalSpend >= HARD_CAP) return { level: 'cap', action: 'stop' };
    if (this.totalSpend >= SOFT_WARN) return { level: 'warn' };
    return { level: 'ok' };
  }
}

```

The system enforces two configurable thresholds: a soft warning and a hard cap. When the cap triggers, the controller automatically invokes `stop()` to terminate the session, ensuring voice interactions never exceed budget constraints.

## Session Lifecycle and Graceful Teardown

Terminating a session requires careful cleanup to release hardware resources and network connections. The `stop()` method in `gevRealtimeController` performs the following sequence:

1. Aborts any pending function calls currently executing in the action runner
2. Closes the `oai-events` data channel to stop event processing
3. Tears down the `RTCPeerConnection` and releases ICE gathering resources
4. Stops all microphone tracks to turn off the recording indicator
5. Resets UI state to idle, clearing visualizers and status messages

This systematic teardown prevents memory leaks and ensures the microphone permission indicator disappears immediately when the user ends the session.

## Code Examples

### Starting a Voice Session

Initialize the controller and begin streaming with push-to-talk disabled for hands-free operation:

```javascript
import { initGevVoiceCommands } from './voice/gevRealtime.js';

const controller = initGevVoiceCommands({ 
  viewer, 
  styleManager, 
  dataManager 
});

// Begin session: fetches token, establishes WebRTC, starts microphone
controller.start({ pushToTalk: false });

```

### Sending Text Commands

While the voice session remains active, inject text-based commands directly into the conversation:

```javascript
controller.sendTextCommand('Show me the traffic in Osaka at 5pm');

```

This method creates a `conversation.item.create` event on the data channel and queues a response generation from the Realtime API.

### Handling Function Calls

The action runner implements domain-specific tools that OpenAI can invoke. For example, radio control:

```javascript
// src/voice/gevActions.js
export async function control_radio({ action }) {
  if (action === 'play') {
    await radioLayer.playback();
    return { 
      ok: true, 
      action, 
      radioPlaybackRequested: true 
    };
  }
  // Handle pause, stop, volume adjustments...
}

```

### Monitoring Session Costs

Access real-time spending data through the cost tracker:

```javascript
const snapshot = controller.costTracker.state();
console.log(`Current spend: ${snapshot.display}`); // "~$0.42"

if (snapshot.level === 'cap') {
  console.warn('Spend cap reached – session stopping');
  controller.stop();
}

```

## Summary

- **Token-based authentication**: The system proxys token requests through `/api/realtime/token` in [`vite.config.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/vite.config.js) to avoid exposing API keys in the client
- **Native WebRTC integration**: `GevRealtimeController` manages `RTCPeerConnection`, SDP exchange with `https://api.openai.com/v1/realtime/calls`, and the `oai-events` data channel
- **Bidirectional audio**: Microphone streams upload via WebRTC while assistant responses play through a standard HTML audio element
- **Event-driven architecture**: The controller handles `input_audio_buffer.speech_started`, `response.done`, and `function_call` events to coordinate the conversation flow
- **Cost safety**: `VoiceCostTracker` in [`voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/voiceCost.js) monitors usage from each response and enforces hard spending caps to prevent budget overruns
- **Clean teardown**: The `stop()` method systematically releases WebRTC resources, data channels, and microphone tracks

## Frequently Asked Questions

### How does the voice system authenticate with the OpenAI Realtime API?

The system uses ephemeral bearer tokens obtained through a backend proxy endpoint. When a session starts, the client requests a token from `/api/realtime/token`, which the development server defined in [`vite.config.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/vite.config.js) generates by forwarding the request to OpenAI with the appropriate model tier. These short-lived tokens authenticate the WebRTC handshake without exposing permanent API credentials in browser code.

### What WebRTC components handle the audio streaming?

The implementation uses a standard `RTCPeerConnection` configured with STUN servers for NAT traversal. The controller adds local microphone tracks to the connection for upstream audio, while the remote audio stream attaches to an HTML `<audio>` element for playback. A dedicated data channel named `oai-events` carries JSON signaling events alongside the media stream.

### How are function calls from the Realtime API executed?

When the controller receives a `function_call` event over the data channel, it extracts the function name and arguments, then dispatches them to the action runner implemented in [`src/voice/gevActions.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevActions.js). After executing the requested operation (such as controlling radio playback or toggling layer visibility), the controller formats the result as a `function_call_output` event and transmits it back to OpenAI to continue the conversation.

### How does the system prevent excessive billing during voice sessions?

A `VoiceCostTracker` instance monitors token usage from every `response.done` event and converts it to USD using model-specific rates. The tracker maintains soft warning and hard cap thresholds. When spending reaches the hard cap, the controller automatically triggers `stop()` to terminate the WebRTC connection and end the session immediately, ensuring costs remain bounded regardless of conversation length.