How God's Eye View Implements Voice Control: A 3-Layer Architecture Deep Dive

God's Eye View implements voice control as a self-contained, three-layer architecture separating UI persistence, session management, and business logic, built on OpenAI's Realtime API with WebRTC data channels and local cost-guardrails.

This open-source geospatial visualization project orchestrates natural language commands for map navigation and radio control through a modular, state-machine-driven system. The voice control architecture demonstrates production-grade patterns for AI-powered interfaces—complete with spend caps, push-to-talk mechanics, and graceful error handling. Let's examine how bilawalsidhu/gods-eye-view structures this subsystem.

The Three-Layer Voice Control Architecture

The voice system is organized into clearly separated concerns, each with dedicated responsibilities and source files:

Layer Responsibility Main Source File
UI & Persistence Renders the on-screen voice widget, stores model tier and cost limits in localStorage src/voice/gevRealtime.js (lines 36-80)
Session Controller Manages Realtime session lifecycle—tokens, WebRTC connection, microphone, push-to-talk, visualization, cost tracking src/voice/gevRealtime.js (lines 90-210)
Business-Logic Runner Executes LLM-issued tool calls for map navigation and radio control src/voice/gevActions.js

This separation enables independent testing, swapping of backing models, and safe evolution of the UI without touching networking or domain logic.

Layer 1: UI and Persistence

The voice widget is constructed by createVoiceControl, invoked from the exported initGevVoiceCommands function. This layer handles:

  • Visual components: activation button, tier selector dropdown, dual audio visualizers, and real-time status read-out
  • Settings persistence: tier selection and spend limits survive page reloads via localStorage

Key persistence functions in src/voice/gevRealtime.js:

// Lines 60-78: Reading and writing user preferences
function readStoredVoiceTier() {
  return localStorage.getItem('godsEyeView.voiceCost.tier') || 'standard';
}

function writeStoredVoiceTier(tier) {
  localStorage.setItem('godsEyeView.voiceCost.tier', tier);
  return tier;
}

function readStoredVoiceLimits() { /* ... */ }
function writeStoredVoiceLimits(limits) { /* ... */ }

The storage keys godsEyeView.voiceCost.tier and godsEyeView.voiceCost.limits namespace settings to avoid collisions with other application data.

Layer 2: Session Controller (GevRealtimeController)

The core orchestration lives in GevRealtimeController, a class managing the full Realtime API lifecycle. This is where voice control architecture complexity concentrates.

Connection Lifecycle (lines 73-99)

The start() method establishes the full pipeline:

async start() {
  // 1. Acquire microphone stream
  this.micStream = await navigator.mediaDevices.getUserMedia({ audio: true });
  
  // 2. Fetch ephemeral token from backend
  const { token, minted } = await fetchRealtimeToken(this.voiceTier);
  
  // 3. Initialize cost tracker with model information
  this.costTracker = createVoiceCostTracker({
    tier: this.voiceTier,
    limits: this.voiceLimits,
    modelInfo: minted.model
  });
  
  // 4. Create WebRTC peer connection and data channel
  this.pc = new RTCPeerConnection({ iceServers: [] });
  this.dc = this.pc.createDataChannel('oai-events');
  
  // 5. Complete SDP handshake (offer → answer)
  const offer = await this.pc.createOffer();
  await this.pc.setLocalDescription(offer);
  // ... signal to OpenAI Realtime endpoint, receive answer
}

Push-to-Talk Implementation (lines 175-210)

Rather than always-listening mode, God's Eye View implements push-to-talk via the Space key:

bindPushToTalkShortcut() {
  window.addEventListener('keydown', (e) => {
    if (e.code === 'Space' && !e.repeat && !this.isRecording) {
      this.startRecording();
    }
  });
  
  window.addEventListener('keyup', (e) => {
    if (e.code === 'Space' && this.isRecording) {
      this.stopRecording();
    }
  });
}

This design choice reduces token consumption and prevents accidental activation during typing.

Audio Visualization

Two independent Web Audio API analyzers drive the UI:

  • startVoiceVisualizer: Processes microphone input for live feedback
  • startAssistantVoiceVisualizer: Renders assistant audio output

Both feed bar-height calculations to the DOM, providing immediate visual confirmation of audio flow in both directions.

Safety: Graceful Teardown (lines 332-368)

The controller guarantees cleanup through fatalError, stop, and connection-state guards:

handleConnectionStateChange(state) {
  if (state === 'failed' || state === 'closed') {
    this.fatalError('Connection lost');
  }
}

stop() {
  // Ensure microphone is always released
  this.micStream?.getTracks().forEach(t => t.stop());
  this.pc?.close();
  this.dc?.close();
  this.syncCostUi(); // Final cost update
}

This pattern prevents the common failure mode of "orphaned microphone access" in WebRTC applications.

Layer 3: Business-Logic Runner (gevActions)

Tool calls from the LLM dispatch through createGevActionRunner, which maps function names to concrete operations. The controller extracts calls from Realtime events and executes them:

// Inside handleRealtimeEvent processing
const calls = extractFunctionCalls(event);
for (const call of deduplicateCalls(calls)) {
  const result = await this.runner(call.name, parsedArguments, {
    signal: this.abortController.signal,
    isCurrent: () => !this.sessionAborted
  });
  // ... send result back to Realtime API
}

Available tools include:

Function Purpose
fly_to_location Animate Cesium camera to coordinates
control_radio Pause, play, or duck radio streams
set_layer_visibility Show/hide map data layers

The runner receives the full application context (viewer, dataManager, etc.) enabling direct manipulation of the 3-D scene.

Cost Control Subsystem (voiceCost.js)

God's Eye View integrates spend protection via VoiceCostTracker in src/voice/voiceCost.js. This demonstrates production-grade operational safety for pay-per-use AI APIs.

Model Registry (lines 49-81)

const VOICE_MODELS = {
  standard: {
    name: 'gpt-4o-realtime-preview',
    inputCostPer1mTokens: 5.00,   // $5 per million
    outputCostPer1mTokens: 20.00  // $20 per million
  },
  mini: {
    name: 'gpt-4o-mini-realtime-preview',
    inputCostPer1mTokens: 0.60,
    outputCostPer1mTokens: 2.40
  }
};

Pricing remains current as of the implementation date; the registry pattern allows easy updates without touching controller code.

Spend Limits (lines 70-73)

const VOICE_COST_LIMITS = {
  softWarningUsd: 2.00,
  hardCapUsd: 5.00
};

The tracker emits three cost levels:

  • "ok": Below soft warning—normal operation
  • "warn": Exceeds $2—UI shows amber indicator
  • "cap": Exceeds $5—session terminates, user must acknowledge

Cost Estimation Logic (lines 69-107)

function estimateUsageCostUsd(usage, tier) {
  const model = VOICE_MODELS[tier];
  const inputCost = (usage.inputTokens / 1_000_000) * model.inputCostPer1mTokens;
  const outputCost = (usage.outputTokens / 1_000_000) * model.outputCostPer1mTokens;
  return inputCost + outputCost;
}

The controller polls accumulated usage from Realtime API events and updates the tracker each cycle.

Initialization and Integration

From the application entry point in src/main.js:

import { initGevVoiceCommands } from './voice/gevRealtime.js';

const voiceController = initGevVoiceCommands({
  viewer,           // Cesium 3-D globe
  styleManager,     // CSS-in-JS helper
  dataManager,      // Layer and radio state
  sceneDirector,    // Optional: scripted sequences
  annotations,      // Optional: markup system
});

// Exposed globally for debugging
window.__gevVoiceCommands = voiceController;

The function returns a fully configured GevRealtimeController instance, ready for user activation.

Tier Switching Dynamics

Users toggle between standard and mini models via the UI:

toggleVoiceTier() {
  const newTier = this.voiceTier === 'standard' ? 'mini' : 'standard';
  this.voiceTier = writeStoredVoiceTier(newTier);
  this.voiceLimits = readStoredVoiceLimits();
  
  // Recreate tracker for new session
  this.costTracker = createVoiceCostTracker({
    tier: this.voiceTier,
    limits: this.voiceLimits
  });
  this.syncCostUi();
}

Important: Tier changes apply to the next session start, not the current connection. This avoids mid-stream model switches that would corrupt the Realtime API state.

Radio Coordination Pattern

Voice sessions and radio playback compete for the audio output device. The controller implements a reservation system:

if (call.name === 'control_radio') {
  // Block concurrent tool calls
  const reservation = this.reserveRadioToolHandoff({ 
    abortScope: 'playback' 
  });
  
  const result = await this.runner(call.name, parsedArguments, {
    signal, 
    isCurrent: reservation.isCurrent
  });
  
  // Defer radio restoration until assistant finishes speaking
  if (result?.ok && result.radioPlaybackRequested) {
    this.pendingRadioPlaybackResult = result;
  }
}

This ensures the assistant's response completes audibly before radio resumes, preventing jarring overlaps.

Key Source Files Reference

File Lines of Responsibility
src/voice/voiceCost.js Model pricing, USD conversion, spend-cap state machine
src/voice/gevRealtime.js Full controller: UI, WebRTC, lifecycle, visualization, cost sync
src/voice/gevActions.js Tool dispatch: map navigation, radio control
src/voice/gevRealtime.test.mjs Unit tests for session, cost-cap, and push-to-talk
src/main.js Application bootstrap and controller injection

Summary

  • Three-layer separation: UI/persistence, session control, and business logic remain independent and testable
  • Production-grade safety: Cost tracker with soft warnings and hard caps prevents runaway API spend
  • Push-to-talk UX: Space-key activation reduces token waste and accidental triggers
  • WebRTC data channels: Direct peer connection to OpenAI Realtime API minimizes latency
  • Graceful degradation: Fatal error handling guarantees microphone release and connection cleanup
  • Radio coordination: Reservation pattern prevents audio output conflicts between voice and background sources

Frequently Asked Questions

What AI model powers the voice control in God's Eye View?

God's Eye View supports two OpenAI Realtime API models: gpt-4o-realtime-preview (standard tier) and gpt-4o-mini-realtime-preview (mini tier). The standard model offers higher capability at $5/$20 per million input/output tokens, while mini reduces cost to $0.60/$2.40 per million tokens. Users toggle between tiers via the UI widget, with selection persisted in localStorage for subsequent sessions.

How does the voice system prevent excessive API costs?

The VoiceCostTracker in src/voice/voiceCost.js implements a two-threshold spend guard. A soft warning triggers at $2 cumulative spend (amber UI indicator), and a hard cap terminates the session at $5, requiring explicit user acknowledgment to resume. All token usage converts to USD estimates in real-time using the VOICE_MODELS registry pricing.

Can the voice commands control features beyond map navigation?

Yes. The gevActions dispatcher supports multiple tool categories: camera movement (fly_to_location), radio playback control (control_radio), and layer visibility toggling. The runner receives full application context, enabling extension for additional domain operations without modifying the core session controller.

Why does push-to-talk use the Space key instead of wake words?

The Space-key implementation in bindPushToTalkShortcut reduces both token consumption and false activations compared to continuous listening. This design choice aligns with the application's operational context—users typically have hands on keyboards during map interaction—and complements the cost-control objectives by minimizing unnecessary audio transmission to the Realtime API.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →