How to Use the OpenAI API for Voice Control in God’s Eye View
God’s Eye View implements voice control by integrating the OpenAI Realtime API into a client-server architecture that streams microphone audio, executes tool calls for in-app actions, and generates spoken HUD summaries via a secure Node.js proxy.
God’s Eye View (GEV) is an open-source geospatial visualization platform that enables hands-free operation through the OpenAI API for voice control. This guide explains the implementation details, configuration steps, and code patterns required to enable realtime voice interaction in your own GEV deployment.
Architecture Overview
The voice system uses a secure client-server split to protect API credentials while maintaining low-latency audio streaming. The browser manages the user interface and audio capture, while the server handles authentication and proxies sensitive requests to OpenAI.
Client-Side Voice Module
The src/voice/gevRealtime.js file initializes a WebSocket connection to OpenAI’s realtime endpoint and manages bidirectional audio streaming. This module handles microphone input, plays back model responses, and dispatches incoming events to the appropriate handlers.
Server-Side Proxy
The server-side implementation in server/providers/openai/realtime.js generates short-lived client secrets and forwards session configuration to OpenAI. This proxy ensures your OPENAI_API_KEY never exposes to the browser. The server also hosts the HUD summary endpoint in server/providers/openai/hud-summary.js, which provides concise five-word intelligence briefings.
Configuring the OpenAI API Integration
Before enabling voice control, you must configure the API credentials and environment variables that authenticate your server with OpenAI’s realtime services.
Environment Variables
Copy .env.example to .env in your project root and add your OpenAI API key. The server also supports storing credentials in the macOS keychain under openai-api/api-key.
# .env
OPENAI_API_KEY=sk-YOUR_OPENAI_KEY
# Optional: limit calls per minute
OPENAI_REALTIMET_TOKEN_RATE_LIMIT=30
Server Initialization
Start the development server to load the environment configuration and initialize the OpenAI middlewares. The server registers endpoints for /api/realtime/token and /api/openai/hud-summary during startup.
npm install
npm run dev
Implementing Realtime Voice Connections
The voice interaction flow follows a three-step process: token generation, WebSocket handshake, and continuous audio streaming.
Requesting a Realtime Token
The browser requests a session token from the /api/realtime/token endpoint. The server creates this token by POSTing the full session configuration—including instructions from server/providers/openai/instructions.js and tool schemas from server/providers/openai/tools.js—to https://api.openai.com/v1/realtime/client_secrets.
import { getRealtimeToken } from './api';
async function initVoice() {
const { token, model, instructions, tools } = await getRealtimeToken();
// Token is a short-lived client secret for the WebSocket
}
Establishing the WebSocket
With the token returned, the client constructs a WebSocket URL pointing to wss://api.openai.com/v1/realtime and appends the model name and authentication parameters.
const ws = new WebSocket(
`wss://api.openai.com/v1/realtime?model=${encodeURIComponent(model)}`,
[token]
);
const gev = new GEVRealtime(ws, { instructions, tools });
gev.start(); // Begins microphone streaming
Streaming Audio and Events
The GEVRealtime class streams PCM audio to the WebSocket and listens for JSON events such as response, tool_call, and error. When the model returns a tool call, the dispatcher routes it to src/voice/gevActions.js for execution.
Handling Voice Commands with Tool Calls
Voice commands materialize as structured tool calls from the OpenAI model. The system maps these calls to specific functions within the God’s Eye View interface.
Mapping Tools to Actions
The src/voice/gevActions.js file translates tool calls into in-app actions like toggling map layers or adjusting camera altitude. The available tools are defined in server/providers/openai/tools.js inside the GEV_REALTIME_TOOLS object, which exports JSON schemas describing each function’s parameters.
When the model emits a tool_call event, gevActions looks up the corresponding implementation, executes the action locally, and returns a result object to the model to continue the conversation.
Creating Custom Voice Tools
Extend the voice capabilities by adding entries to server/providers/openai/tools.js and implementing handlers in src/voice/gevActions.js.
// server/providers/openai/tools.js
export const GEV_REALTIME_TOOLS = {
...existingTools,
focusLayer: {
description: "Focus the camera on a specific map layer",
parameters: {
type: "object",
properties: {
layerId: { type: "string" }
},
required: ["layerId"]
}
}
};
// src/voice/gevActions.js
export async function handleToolCall(call) {
switch (call.name) {
case 'focusLayer':
await focusMapLayer(call.parameters.layerId);
return { result: `Focused on layer ${call.parameters.layerId}` };
default:
throw new Error(`Unknown tool: ${call.name}`);
}
}
Generating HUD Summaries
The system provides concise situational awareness through the /api/openai/hud-summary endpoint implemented in server/providers/openai/hud-summary.js. When a user requests a summary via voice, the client queries this endpoint, which forwards a prompt to the gpt‑5‑nano model (default configuration) and returns a terse five-word response.
The action dispatcher in src/voice/gevActions.js handles the hudSummary intent by fetching from this endpoint and returning the text to be spoken back to the pilot.
case 'hudSummary':
const resp = await fetch('/api/openai/hud-summary');
const { summary } = await resp.json();
return { result: summary };
Summary
- Secure Authentication: The server generates short-lived tokens via
server/providers/openai/realtime.jsto protect yourOPENAI_API_KEYfrom client exposure. - Realtime Streaming:
src/voice/gevRealtime.jsmanages the WebSocket connection towss://api.openai.com/v1/realtimeand handles bidirectional audio streaming. - Tool Integration: Voice commands map to JavaScript functions through
src/voice/gevActions.jsusing schemas defined inserver/providers/openai/tools.js. - HUD Summaries: The
/api/openai/hud-summaryendpoint provides spoken intelligence briefings powered by OpenAI language models.
Frequently Asked Questions
What OpenAI model does God’s Eye View use for voice control?
God’s Eye View connects to OpenAI’s Realtime API for voice interactions, using the realtime models for bidirectional audio. For HUD summaries, the default configuration in server/providers/openai/hud-summary.js uses gpt‑5‑nano, though this can be modified in the server configuration.
How does God’s Eye View secure the OpenAI API key?
The system never exposes the OPENAI_API_KEY to the browser. Instead, the server in server/providers/openai/realtime.js creates ephemeral client secrets via the OpenAI API and passes only these short-lived tokens to the WebSocket client, maintaining secure credential isolation.
Can I add custom voice commands to God’s Eye View?
Yes, extend GEV_REALTIME_TOOLS in server/providers/openai/tools.js with your JSON schema, then add the corresponding handler logic in src/voice/gevActions.js. The model will automatically recognize and invoke your custom tools during voice conversations.
What audio format does the realtime API use?
The src/voice/gevRealtime.js module streams PCM audio to the WebSocket connection. The implementation handles the encoding and decoding required by the OpenAI Realtime API, allowing you to work with standard web audio APIs on the client side.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →