What API Is Used for Voice Control in God's Eye View?

God's Eye View implements voice control using the OpenAI Realtime API, establishing a WebSocket connection to the GPT Realtime voice endpoint at https://api.openai.com/v1/realtime/calls.

The open-source repository bilawalsidhu/gods-eye-view provides a hands-free interface driven by OpenAI's Realtime voice models. Understanding the underlying API is critical for developers who need to customize voice tiers, manage costs, or debug connection issues.

OpenAI Realtime API Implementation

The voice subsystem relies on the OpenAI Realtime API to stream audio bidirectionally with low latency. Rather than using traditional request-response HTTP, the application opens a persistent WebSocket to OpenAI's Realtime Calls service, enabling continuous conversation flow.

Configuration and Model Resolution

Model selection and pricing logic reside in src/voice/voiceCost.js. This file exports the resolveVoiceModel() function, which maps the OPENAI_REALTIME_VOICE environment variable to specific OpenAI model identifiers, defaulting to the "marin" tier when no override is provided. The file also catalogs available voice tiers and their associated costs according to OpenAI's Realtime pricing documentation.

WebSocket Client Logic

The actual network transport is implemented in src/voice/gevRealtime.js. This module defines the constant REALTIME_CALLS_URL = 'https://api.openai.com/v1/realtime/calls' and manages the WebSocket handshake using the realtime subprotocol. It handles authentication token injection, message framing, and connection state management for the audio stream.

Configuring the Voice Model

Before starting the application, define your preferred voice tier in the environment. The build system (referenced in vite.config.js) injects this value into the client bundle.


# .env

OPENAI_REALTIME_VOICE=marin

Available values correspond to the tiers defined in src/voice/voiceCost.js, such as marin or other supported OpenAI Realtime models.

Establishing the Realtime Connection

The following pattern demonstrates how the application initializes the WebSocket using the URL constant and model resolver from the voice modules:

import { REALTIME_CALLS_URL } from './src/voice/gevRealtime.js';
import { resolveVoiceModel } from './src/voice/voiceCost.js';

// Resolve the model identifier (defaults to "marin")
const voiceModel = resolveVoiceModel(
  process.env.OPENAI_REALTIME_VOICE || 'marin'
);

// Construct the authenticated WebSocket URL
const socketUrl = `${REALTIME_CALLS_URL}?model=${voiceModel}`;

// Open the connection with the 'realtime' subprotocol
const ws = new WebSocket(socketUrl, ['realtime']);

ws.onopen = () => console.log('Voice engine connected');
ws.onmessage = (event) => console.log('Model response:', event.data);
ws.onerror = (err) => console.error('Realtime connection error:', err);

Streaming Audio to the Model

Once the WebSocket is open, the client streams raw audio data as JSON messages. The payload specifies the message type and includes the binary audio chunk, typically as a base64-encoded string within the JSON frame.

// audioChunk should be an ArrayBuffer or base64 string of recorded audio
function streamAudioChunk(audioChunk) {
  const payload = {
    type: 'input_audio',
    audio: audioChunk,
    // Additional metadata such as sample rate can be included here
  };
  
  if (ws.readyState === WebSocket.OPEN) {
    ws.send(JSON.stringify(payload));
  }
}

Key Source Files

Summary

  • God's Eye View uses the OpenAI Realtime API for voice control, connecting to https://api.openai.com/v1/realtime/calls.
  • The src/voice/gevRealtime.js file manages the WebSocket transport and defines the REALTIME_CALLS_URL endpoint.
  • Model resolution and pricing are handled in src/voice/voiceCost.js via the resolveVoiceModel() function.
  • Set the OPENAI_REALTIME_VOICE environment variable to override the default "marin" model tier.

Frequently Asked Questions

What API powers the voice control in God's Eye View?

The system uses the OpenAI Realtime API, specifically the GPT Realtime voice endpoint that supports streaming audio over WebSockets. This enables natural, low-latency conversation without polling.

How do I change the voice model in God's Eye View?

Set the OPENAI_REALTIME_VOICE environment variable to your desired tier (e.g., marin) before building or running the application. The resolveVoiceModel() function in src/voice/voiceCost.js processes this value and validates it against available OpenAI models.

Where is the WebSocket connection logic implemented?

The core WebSocket client logic resides in src/voice/gevRealtime.js. This file exports the REALTIME_CALLS_URL constant and manages the connection lifecycle, including the realtime subprotocol handshake required by OpenAI.

What is the default voice model if none is specified?

If the OPENAI_REALTIME_VOICE environment variable is absent, the application defaults to the "marin" model tier. This fallback is defined in the resolveVoiceModel() utility within src/voice/voiceCost.js.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →