Voice Control Requirements in God's Eye View: Complete Setup Guide
To enable voice control in God's Eye View, you must provide an OpenAI API key through environment variables or the in-app provider settings, grant browser microphone permissions for the GEV MIC interface, and optionally configure model tiers and spend-guard limits to manage costs.
God's Eye View is an open-source geospatial visualization project that integrates OpenAI's Realtime API to deliver hands-free voice interaction. Understanding the voice control requirements ensures you can leverage the GEV MIC interface without configuration errors. The implementation centers on a minimal set of secrets and permissions defined in the bilawalsidhu/gods-eye-view repository.
Core Requirements for Voice Control
OpenAI API Key Configuration
The only mandatory secret for voice functionality is the OpenAI API key. You can set OPENAI_API_KEY in a local .env file based on the provided .env.example, or add it later through the POWER UP → Provider Settings panel in the UI. Without this key, the mic button appears but remains disabled, preventing all voice operations.
Browser Microphone Permissions
The UI element GEV MIC requests access to the microphone through standard browser APIs. When prompted, users must grant permission; otherwise, the voice control system cannot initialize. This logic is managed in src/voice/gevRealtime.js, which coordinates the OpenAI Realtime session with the browser's media streams.
Optional Configuration Settings
Model Tier Selection
You can choose between standard and cost-optimized "mini" models by overriding environment variables before startup:
OPENAI_REALTIME_MODEL(default:gpt-realtime-2)OPENAI_REALTIME_MODEL_MINI(default:gpt-realtime-2.1-mini)
These values control which model identifier the cost tracker uses and which endpoint the Realtime session targets.
Spend-Guard Limits
The client enforces cost controls defined in src/voice/voiceCost.js. The VOICE_COST_LIMITS object sets a soft warning at ~$2 and a hard cap at ~$5 per session. The UI displays real-time costs (~$0.xx) and automatically terminates the session when the cap is reached, preventing unexpected API charges.
Implementation Architecture
Voice Cost Tracking
The createVoiceCostTracker function in src/voice/voiceCost.js initializes a cost monitor with configurable limits. This tracker records usage after each Realtime response to enforce spend guards programmatically.
import { createVoiceCostTracker } from './voice/voiceCost.js';
// Initialize with custom limits
const tracker = createVoiceCostTracker({
modelId: process.env.OPENAI_REALTIME_MODEL,
limits: { warnUsd: 2, capUsd: 5 }
});
// Record usage after each response
tracker.record(response.usage);
Realtime Session Management
src/voice/gevRealtime.js handles the WebRTC connection to OpenAI's Realtime API, manages the GEV MIC button state, and processes audio streams. It relies on the system prompt defined in server/providers/openai/instructions.js, which configures the agent behavior ("You are GEV Voice Control...").
Quick Start Configuration
Create a .env file in the project root to satisfy voice control requirements:
# .env
OPENAI_API_KEY=sk-your-openai-key-here
# Optional: Model selection
OPENAI_REALTIME_MODEL=gpt-realtime-2
OPENAI_REALTIME_MODEL_MINI=gpt-realtime-2.1-mini
For runtime configuration without file editing:
- Start the application with
npm run dev - Click the GEV MIC button in the lower-right corner
- Allow microphone access when the browser prompts
- Issue commands such as "Take me to LAX and select the nearest airborne aircraft"
Summary
- OpenAI API Key: Set via
OPENAI_API_KEYenvironment variable or the POWER UP → Provider Settings panel; the only required secret for voice. - Microphone Permission: Grant browser access when prompted by the GEV MIC interface controlled by
gevRealtime.js. - Optional Tiers: Configure
OPENAI_REALTIME_MODELandOPENAI_REALTIME_MODEL_MINIin.envto switch between standard and mini models. - Cost Controls: Built-in spend guards in
voiceCost.jswarn at $2 and hard-cap at $5 per session. - Independence: Voice control requires no map API keys (Cesium Ion or Google Maps), though those are needed for imagery layers.
Frequently Asked Questions
What happens if I don't set an OpenAI API key?
Without the OPENAI_API_KEY environment variable or in-app configuration, the GEV MIC button remains visible but disabled. Voice control functions are inoperative, though the rest of the application continues to work normally.
Can I use voice control without a Cesium Ion token?
Yes. Voice control depends solely on the OpenAI Realtime API and browser microphone access. Cesium Ion tokens and Google Maps keys are only required for map imagery layers and do not affect voice functionality.
How do I switch to the cheaper mini model for voice?
Set OPENAI_REALTIME_MODEL_MINI=gpt-realtime-2.1-mini in your .env file before starting the application. The system uses this identifier when initializing cost tracking in voiceCost.js and establishing the Realtime session.
What happens when I reach the $5 spend limit?
Once the hard cap defined in VOICE_COST_LIMITS is reached, the voice session automatically terminates. The UI displays the final cost and prevents further Realtime API calls until a new session begins, protecting against runaway API expenses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →