How Voice Commands Are Translated into Actions Using the GEV Action Runner
The GEV action runner—implemented in src/voice/gevActions.js—receives structured tool calls from the LLM, resolves human-friendly aliases to internal IDs, normalizes arguments, and dispatches commands to Cesium and UI services to execute the requested action.
God’s Eye View (GEV) provides a voice-control pipeline that converts natural speech into concrete application behavior. The translation layer centers on two modules: gevRealtime.js, which manages audio capture and LLM communication, and gevActions.js, which implements the action runner responsible for turning LLM outputs into camera movements, UI updates, and data layer changes.
The Voice Control Architecture
Real-time Processing in gevRealtime.js
Voice interaction begins in src/voice/gevRealtime.js inside the GevRealtimeController class. This controller captures the microphone stream, sends the audio to a large language model (LLM), and awaits a structured tool call response. When the LLM returns a JSON object containing the intended action—such as {name: "adjust_camera_zoom", arguments: {direction: "in"}}—the controller immediately forwards it to the action runner.
At approximately line 1215, the controller executes the runner with the parsed arguments:
// gevRealtime.js – line ~1215
result = await this.runner(call.name, parsedArguments, {
// context: viewer, dataManager, current UI state, etc.
});
The controller also initializes the runner during setup (around line 200) by injecting required services:
// gevRealtime.js – line ~200
const runner = createGevActionRunner({
viewer,
styleManager,
dataManager,
sceneDirector,
annotations
});
The Tool Call Interface
The contract between the LLM and the action runner is strictly typed. The LLM emits a tool call containing a name string and an arguments object. The runner validates these inputs, maps them to internal vocabulary, and triggers the corresponding handler. This decoupling ensures that natural language variations like “zoom in” or “get closer” both resolve to the adjust_camera_zoom action with normalized parameters.
How the Action Runner Processes Commands
The createGevActionRunner factory function in src/voice/gevActions.js returns the core dispatch function that handles all voice-initiated operations. The runner performs four distinct phases before executing side effects.
Initialization with createGevActionRunner
The runner requires access to the Cesium viewer, data managers, and UI controllers to perform its work. When created, it closures over these dependencies, making them available to every subsequent command without requiring global state. This pattern keeps the voice layer testable and isolated from the main application logic.
Normalizing Input with Alias Resolution
Human-friendly terms often differ from internal panel IDs. Lines 25–44 of gevActions.js define a PANEL_ALIASES Map that translates colloquial names to machine identifiers:
// gevActions.js – lines 25-44
const PANEL_ALIASES = new Map([
['data', 'data-panel'],
['cctv', 'cctv-panel'],
['layers', 'data-panel'],
// additional mappings...
]);
When the LLM refers to “cctv” or “data layers,” the runner resolves these strings to the canonical panel IDs before attempting to toggle visibility or focus.
Context Mode Vocabulary Translation
Certain commands include a mode field that requires semantic translation. The withContextModeVocabulary function (lines 70–124) rewrites LLM-specific mode strings into the public vocabulary expected by the UI state manager. This ensures that phrases like “switch to tactical view” map to the correct internal enumeration regardless of how the LLM phrases the mode name.
Command Dispatch Logic
The core of the runner is a large conditional chain spanning roughly lines 317–980 that matches the normalized action string to concrete handlers. Each branch constructs a result object confirming the operation and its parameters. For example, handling a zoom request:
// gevActions.js – around line 317
if (action === 'adjust_camera_zoom') {
const zoomOut = args.direction === 'out';
const amt = Math.abs(args.amount || 1);
// ...Cesium camera logic...
return {
ok: true,
action: 'adjust_camera_zoom',
direction: zoomOut ? 'out' : 'in',
amount: amt,
orbitRadiusAdjusted: true
};
}
Each handler returns a structured result that gevRealtime.js consumes to provide verbal confirmation and trigger follow-up operations like pausing radio playback during camera transitions.
Complete Execution Flow Example
The following sequence illustrates how a spoken command flows through the system:
-
User speaks “Zoom in on the target” into the microphone.
-
GevRealtimeControllerstreams audio to the LLM and receives:const toolCall = { name: 'adjust_camera_zoom', arguments: { direction: 'in', amount: 2 } }; -
Controller invokes the runner:
const result = await runner(toolCall.name, toolCall.arguments, context); -
Action runner resolves aliases, validates arguments, adjusts the Cesium camera via the
viewerinstance, and returns:{ ok: true, action: 'adjust_camera_zoom', direction: 'in', amount: 2, orbitRadiusAdjusted: true } -
Controller provides audio feedback—“Zooming in”—and updates any relevant UI indicators.
This pipeline ensures that voice commands translated into actions are deterministic, auditable, and cleanly separated from the speech recognition and synthesis layers.
Summary
- Entry Point:
src/voice/gevRealtime.jscaptures voice input and receives structured tool calls from the LLM. - Translation Layer: The
createGevActionRunnerfactory insrc/voice/gevActions.jsinstantiates the dispatcher with injected Cesium and UI services. - Normalization: The runner uses
PANEL_ALIASESandwithContextModeVocabularyto map human-friendly terms to internal system identifiers. - Execution: A comprehensive conditional chain (lines 317–980) routes commands to specific handlers that perform camera moves, layer toggles, and cockpit controls.
- Feedback: Results bubble back to the controller, which manages side effects and verbal confirmations to complete the interaction loop.
Frequently Asked Questions
What is the role of gevActions.js in the voice pipeline?
gevActions.js implements the createGevActionRunner factory function and contains the full catalogue of executable verbs. It resolves aliases, normalizes arguments, and executes the actual Cesium and UI side effects, acting as the bridge between LLM intent and application behavior.
How does the action runner handle ambiguous voice commands?
The runner relies on the LLM to disambiguate intent before invocation. Once the LLM returns a specific tool call name (e.g., adjust_camera_zoom rather than move_camera), the runner performs strict validation. If arguments are missing or aliases unmapped, the function returns an error result that the controller converts into a clarification prompt for the user.
What services does the action runner interact with?
According to the initialization code in gevRealtime.js line 200, the runner receives the Cesium viewer, styleManager, dataManager, sceneDirector, and annotations services. Handlers in gevActions.js use these to manipulate the 3D globe, toggle data layers, and manage UI state.
How is the runner initialized in the voice controller?
The GevRealtimeController constructs the runner during its setup phase by calling createGevActionRunner with an options object containing required service references. This occurs around line 200 of gevRealtime.js, ensuring the runner is available before any audio processing begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →