What API Is Used for Voice Control in God's Eye View?
God's Eye View implements voice control using the OpenAI Realtime API, establishing a WebSocket connection to the GPT Realtime voice endpoint at https://api.openai.com/v1/realtime/calls.
The open-source repository bilawalsidhu/gods-eye-view provides a hands-free interface driven by OpenAI's Realtime voice models. Understanding the underlying API is critical for developers who need to customize voice tiers, manage costs, or debug connection issues.
OpenAI Realtime API Implementation
The voice subsystem relies on the OpenAI Realtime API to stream audio bidirectionally with low latency. Rather than using traditional request-response HTTP, the application opens a persistent WebSocket to OpenAI's Realtime Calls service, enabling continuous conversation flow.
Configuration and Model Resolution
Model selection and pricing logic reside in src/voice/voiceCost.js. This file exports the resolveVoiceModel() function, which maps the OPENAI_REALTIME_VOICE environment variable to specific OpenAI model identifiers, defaulting to the "marin" tier when no override is provided. The file also catalogs available voice tiers and their associated costs according to OpenAI's Realtime pricing documentation.
WebSocket Client Logic
The actual network transport is implemented in src/voice/gevRealtime.js. This module defines the constant REALTIME_CALLS_URL = 'https://api.openai.com/v1/realtime/calls' and manages the WebSocket handshake using the realtime subprotocol. It handles authentication token injection, message framing, and connection state management for the audio stream.
Configuring the Voice Model
Before starting the application, define your preferred voice tier in the environment. The build system (referenced in vite.config.js) injects this value into the client bundle.
# .env
OPENAI_REALTIME_VOICE=marin
Available values correspond to the tiers defined in src/voice/voiceCost.js, such as marin or other supported OpenAI Realtime models.
Establishing the Realtime Connection
The following pattern demonstrates how the application initializes the WebSocket using the URL constant and model resolver from the voice modules:
import { REALTIME_CALLS_URL } from './src/voice/gevRealtime.js';
import { resolveVoiceModel } from './src/voice/voiceCost.js';
// Resolve the model identifier (defaults to "marin")
const voiceModel = resolveVoiceModel(
process.env.OPENAI_REALTIME_VOICE || 'marin'
);
// Construct the authenticated WebSocket URL
const socketUrl = `${REALTIME_CALLS_URL}?model=${voiceModel}`;
// Open the connection with the 'realtime' subprotocol
const ws = new WebSocket(socketUrl, ['realtime']);
ws.onopen = () => console.log('Voice engine connected');
ws.onmessage = (event) => console.log('Model response:', event.data);
ws.onerror = (err) => console.error('Realtime connection error:', err);
Streaming Audio to the Model
Once the WebSocket is open, the client streams raw audio data as JSON messages. The payload specifies the message type and includes the binary audio chunk, typically as a base64-encoded string within the JSON frame.
// audioChunk should be an ArrayBuffer or base64 string of recorded audio
function streamAudioChunk(audioChunk) {
const payload = {
type: 'input_audio',
audio: audioChunk,
// Additional metadata such as sample rate can be included here
};
if (ws.readyState === WebSocket.OPEN) {
ws.send(JSON.stringify(payload));
}
}
Key Source Files
src/voice/voiceCost.js— Defines voice tier pricing, available models, and theresolveVoiceModel()utility. (View on GitHub)src/voice/gevRealtime.js— ContainsREALTIME_CALLS_URLand the WebSocket client implementation that communicates with OpenAI. (View on GitHub)vite.config.js— Handles build-time injection of theOPENAI_REALTIME_VOICEenvironment variable. (View on GitHub).env.example— Documents the required environment variables for voice configuration. (View on GitHub)
Summary
- God's Eye View uses the OpenAI Realtime API for voice control, connecting to
https://api.openai.com/v1/realtime/calls. - The
src/voice/gevRealtime.jsfile manages the WebSocket transport and defines theREALTIME_CALLS_URLendpoint. - Model resolution and pricing are handled in
src/voice/voiceCost.jsvia theresolveVoiceModel()function. - Set the
OPENAI_REALTIME_VOICEenvironment variable to override the default "marin" model tier.
Frequently Asked Questions
What API powers the voice control in God's Eye View?
The system uses the OpenAI Realtime API, specifically the GPT Realtime voice endpoint that supports streaming audio over WebSockets. This enables natural, low-latency conversation without polling.
How do I change the voice model in God's Eye View?
Set the OPENAI_REALTIME_VOICE environment variable to your desired tier (e.g., marin) before building or running the application. The resolveVoiceModel() function in src/voice/voiceCost.js processes this value and validates it against available OpenAI models.
Where is the WebSocket connection logic implemented?
The core WebSocket client logic resides in src/voice/gevRealtime.js. This file exports the REALTIME_CALLS_URL constant and manages the connection lifecycle, including the realtime subprotocol handshake required by OpenAI.
What is the default voice model if none is specified?
If the OPENAI_REALTIME_VOICE environment variable is absent, the application defaults to the "marin" model tier. This fallback is defined in the resolveVoiceModel() utility within src/voice/voiceCost.js.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →