# How God's Eye View Implements Voice Control: A 3-Layer Architecture Deep Dive

> Explore the 3-layer architecture of God's Eye View voice control. Learn how it uses OpenAI Realtime API, WebRTC, and cost-guardrails for seamless interaction.

- Repository: [Bilawal Sidhu/gods-eye-view](https://github.com/bilawalsidhu/gods-eye-view)
- Tags: architecture
- Published: 2026-09-06

---

**God's Eye View implements voice control as a self-contained, three-layer architecture separating UI persistence, session management, and business logic, built on OpenAI's Realtime API with WebRTC data channels and local cost-guardrails.**

This open-source geospatial visualization project orchestrates natural language commands for map navigation and radio control through a modular, state-machine-driven system. The voice control architecture demonstrates production-grade patterns for AI-powered interfaces—complete with spend caps, push-to-talk mechanics, and graceful error handling. Let's examine how bilawalsidhu/gods-eye-view structures this subsystem.

## The Three-Layer Voice Control Architecture

The voice system is organized into clearly separated concerns, each with dedicated responsibilities and source files:

| Layer | Responsibility | Main Source File |
|-------|----------------|------------------|
| **UI & Persistence** | Renders the on-screen voice widget, stores model tier and cost limits in `localStorage` | [`src/voice/gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevRealtime.js) (lines 36-80) |
| **Session Controller** | Manages Realtime session lifecycle—tokens, WebRTC connection, microphone, push-to-talk, visualization, cost tracking | [`src/voice/gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevRealtime.js) (lines 90-210) |
| **Business-Logic Runner** | Executes LLM-issued tool calls for map navigation and radio control | [`src/voice/gevActions.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevActions.js) |

This separation enables independent testing, swapping of backing models, and safe evolution of the UI without touching networking or domain logic.

## Layer 1: UI and Persistence

The voice widget is constructed by **`createVoiceControl`**, invoked from the exported **`initGevVoiceCommands`** function. This layer handles:

- **Visual components**: activation button, tier selector dropdown, dual audio visualizers, and real-time status read-out
- **Settings persistence**: tier selection and spend limits survive page reloads via `localStorage`

Key persistence functions in [`src/voice/gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevRealtime.js):

```javascript
// Lines 60-78: Reading and writing user preferences
function readStoredVoiceTier() {
  return localStorage.getItem('godsEyeView.voiceCost.tier') || 'standard';
}

function writeStoredVoiceTier(tier) {
  localStorage.setItem('godsEyeView.voiceCost.tier', tier);
  return tier;
}

function readStoredVoiceLimits() { /* ... */ }
function writeStoredVoiceLimits(limits) { /* ... */ }

```

The storage keys `godsEyeView.voiceCost.tier` and `godsEyeView.voiceCost.limits` namespace settings to avoid collisions with other application data.

## Layer 2: Session Controller (`GevRealtimeController`)

The core orchestration lives in **`GevRealtimeController`**, a class managing the full Realtime API lifecycle. This is where voice control architecture complexity concentrates.

### Connection Lifecycle (lines 73-99)

The `start()` method establishes the full pipeline:

```javascript
async start() {
  // 1. Acquire microphone stream
  this.micStream = await navigator.mediaDevices.getUserMedia({ audio: true });
  
  // 2. Fetch ephemeral token from backend
  const { token, minted } = await fetchRealtimeToken(this.voiceTier);
  
  // 3. Initialize cost tracker with model information
  this.costTracker = createVoiceCostTracker({
    tier: this.voiceTier,
    limits: this.voiceLimits,
    modelInfo: minted.model
  });
  
  // 4. Create WebRTC peer connection and data channel
  this.pc = new RTCPeerConnection({ iceServers: [] });
  this.dc = this.pc.createDataChannel('oai-events');
  
  // 5. Complete SDP handshake (offer → answer)
  const offer = await this.pc.createOffer();
  await this.pc.setLocalDescription(offer);
  // ... signal to OpenAI Realtime endpoint, receive answer
}

```

### Push-to-Talk Implementation (lines 175-210)

Rather than always-listening mode, God's Eye View implements **push-to-talk** via the Space key:

```javascript
bindPushToTalkShortcut() {
  window.addEventListener('keydown', (e) => {
    if (e.code === 'Space' && !e.repeat && !this.isRecording) {
      this.startRecording();
    }
  });
  
  window.addEventListener('keyup', (e) => {
    if (e.code === 'Space' && this.isRecording) {
      this.stopRecording();
    }
  });
}

```

This design choice reduces token consumption and prevents accidental activation during typing.

### Audio Visualization

Two independent Web Audio API analyzers drive the UI:

- **`startVoiceVisualizer`**: Processes microphone input for live feedback
- **`startAssistantVoiceVisualizer`**: Renders assistant audio output

Both feed bar-height calculations to the DOM, providing immediate visual confirmation of audio flow in both directions.

### Safety: Graceful Teardown (lines 332-368)

The controller guarantees cleanup through **`fatalError`**, **`stop`**, and connection-state guards:

```javascript
handleConnectionStateChange(state) {
  if (state === 'failed' || state === 'closed') {
    this.fatalError('Connection lost');
  }
}

stop() {
  // Ensure microphone is always released
  this.micStream?.getTracks().forEach(t => t.stop());
  this.pc?.close();
  this.dc?.close();
  this.syncCostUi(); // Final cost update
}

```

This pattern prevents the common failure mode of "orphaned microphone access" in WebRTC applications.

## Layer 3: Business-Logic Runner (`gevActions`)

Tool calls from the LLM dispatch through **`createGevActionRunner`**, which maps function names to concrete operations. The controller extracts calls from Realtime events and executes them:

```javascript
// Inside handleRealtimeEvent processing
const calls = extractFunctionCalls(event);
for (const call of deduplicateCalls(calls)) {
  const result = await this.runner(call.name, parsedArguments, {
    signal: this.abortController.signal,
    isCurrent: () => !this.sessionAborted
  });
  // ... send result back to Realtime API
}

```

Available tools include:

| Function | Purpose |
|----------|---------|
| `fly_to_location` | Animate Cesium camera to coordinates |
| `control_radio` | Pause, play, or duck radio streams |
| `set_layer_visibility` | Show/hide map data layers |

The runner receives the full application context (`viewer`, `dataManager`, etc.) enabling direct manipulation of the 3-D scene.

## Cost Control Subsystem ([`voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/voiceCost.js))

God's Eye View integrates spend protection via **`VoiceCostTracker`** in [`src/voice/voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/voiceCost.js). This demonstrates production-grade operational safety for pay-per-use AI APIs.

### Model Registry (lines 49-81)

```javascript
const VOICE_MODELS = {
  standard: {
    name: 'gpt-4o-realtime-preview',
    inputCostPer1mTokens: 5.00,   // $5 per million
    outputCostPer1mTokens: 20.00  // $20 per million
  },
  mini: {
    name: 'gpt-4o-mini-realtime-preview',
    inputCostPer1mTokens: 0.60,
    outputCostPer1mTokens: 2.40
  }
};

```

Pricing remains current as of the implementation date; the registry pattern allows easy updates without touching controller code.

### Spend Limits (lines 70-73)

```javascript
const VOICE_COST_LIMITS = {
  softWarningUsd: 2.00,
  hardCapUsd: 5.00
};

```

The tracker emits three **cost levels**:

- **`"ok"`**: Below soft warning—normal operation
- **`"warn"`**: Exceeds $2—UI shows amber indicator
- **`"cap"`**: Exceeds $5—session terminates, user must acknowledge

### Cost Estimation Logic (lines 69-107)

```javascript
function estimateUsageCostUsd(usage, tier) {
  const model = VOICE_MODELS[tier];
  const inputCost = (usage.inputTokens / 1_000_000) * model.inputCostPer1mTokens;
  const outputCost = (usage.outputTokens / 1_000_000) * model.outputCostPer1mTokens;
  return inputCost + outputCost;
}

```

The controller polls accumulated usage from Realtime API events and updates the tracker each cycle.

## Initialization and Integration

From the application entry point in [`src/main.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/main.js):

```javascript
import { initGevVoiceCommands } from './voice/gevRealtime.js';

const voiceController = initGevVoiceCommands({
  viewer,           // Cesium 3-D globe
  styleManager,     // CSS-in-JS helper
  dataManager,      // Layer and radio state
  sceneDirector,    // Optional: scripted sequences
  annotations,      // Optional: markup system
});

// Exposed globally for debugging
window.__gevVoiceCommands = voiceController;

```

The function returns a fully configured `GevRealtimeController` instance, ready for user activation.

## Tier Switching Dynamics

Users toggle between standard and mini models via the UI:

```javascript
toggleVoiceTier() {
  const newTier = this.voiceTier === 'standard' ? 'mini' : 'standard';
  this.voiceTier = writeStoredVoiceTier(newTier);
  this.voiceLimits = readStoredVoiceLimits();
  
  // Recreate tracker for new session
  this.costTracker = createVoiceCostTracker({
    tier: this.voiceTier,
    limits: this.voiceLimits
  });
  this.syncCostUi();
}

```

**Important**: Tier changes apply to the *next* session start, not the current connection. This avoids mid-stream model switches that would corrupt the Realtime API state.

## Radio Coordination Pattern

Voice sessions and radio playback compete for the audio output device. The controller implements a **reservation system**:

```javascript
if (call.name === 'control_radio') {
  // Block concurrent tool calls
  const reservation = this.reserveRadioToolHandoff({ 
    abortScope: 'playback' 
  });
  
  const result = await this.runner(call.name, parsedArguments, {
    signal, 
    isCurrent: reservation.isCurrent
  });
  
  // Defer radio restoration until assistant finishes speaking
  if (result?.ok && result.radioPlaybackRequested) {
    this.pendingRadioPlaybackResult = result;
  }
}

```

This ensures the assistant's response completes audibly before radio resumes, preventing jarring overlaps.

## Key Source Files Reference

| File | Lines of Responsibility |
|------|------------------------|
| [`src/voice/voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/voiceCost.js) | Model pricing, USD conversion, spend-cap state machine |
| [`src/voice/gevRealtime.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevRealtime.js) | Full controller: UI, WebRTC, lifecycle, visualization, cost sync |
| [`src/voice/gevActions.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/gevActions.js) | Tool dispatch: map navigation, radio control |
| `src/voice/gevRealtime.test.mjs` | Unit tests for session, cost-cap, and push-to-talk |
| [`src/main.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/main.js) | Application bootstrap and controller injection |

## Summary

- **Three-layer separation**: UI/persistence, session control, and business logic remain independent and testable
- **Production-grade safety**: Cost tracker with soft warnings and hard caps prevents runaway API spend
- **Push-to-talk UX**: Space-key activation reduces token waste and accidental triggers
- **WebRTC data channels**: Direct peer connection to OpenAI Realtime API minimizes latency
- **Graceful degradation**: Fatal error handling guarantees microphone release and connection cleanup
- **Radio coordination**: Reservation pattern prevents audio output conflicts between voice and background sources

## Frequently Asked Questions

### What AI model powers the voice control in God's Eye View?

God's Eye View supports two OpenAI Realtime API models: **gpt-4o-realtime-preview** (standard tier) and **gpt-4o-mini-realtime-preview** (mini tier). The standard model offers higher capability at $5/$20 per million input/output tokens, while mini reduces cost to $0.60/$2.40 per million tokens. Users toggle between tiers via the UI widget, with selection persisted in `localStorage` for subsequent sessions.

### How does the voice system prevent excessive API costs?

The **`VoiceCostTracker`** in [`src/voice/voiceCost.js`](https://github.com/bilawalsidhu/gods-eye-view/blob/main/src/voice/voiceCost.js) implements a two-threshold spend guard. A soft warning triggers at $2 cumulative spend (amber UI indicator), and a hard cap terminates the session at $5, requiring explicit user acknowledgment to resume. All token usage converts to USD estimates in real-time using the `VOICE_MODELS` registry pricing.

### Can the voice commands control features beyond map navigation?

Yes. The **`gevActions`** dispatcher supports multiple tool categories: camera movement (`fly_to_location`), radio playback control (`control_radio`), and layer visibility toggling. The runner receives full application context, enabling extension for additional domain operations without modifying the core session controller.

### Why does push-to-talk use the Space key instead of wake words?

The Space-key implementation in `bindPushToTalkShortcut` reduces both token consumption and false activations compared to continuous listening. This design choice aligns with the application's operational context—users typically have hands on keyboards during map interaction—and complements the cost-control objectives by minimizing unnecessary audio transmission to the Realtime API.