How OpenWhispr Integrates with the MediaRecorder API for Audio Capture

OpenWhispr captures microphone audio using the browser's MediaRecorder API in the renderer process, coordinating between React hooks, the AudioManager helper, and Electron's IPC layer to stream audio chunks to the main process for local or cloud-based transcription.

OpenWhispr is an open-source, privacy-first dictation application built with Electron and React. The application leverages the MediaRecorder API to capture raw audio streams directly in the browser renderer, ensuring sensitive voice data never leaves the device unless a cloud provider is explicitly selected. This integration creates a seamless pipeline from microphone input to speech-to-text transcription through a tightly coupled system of hooks, helpers, and inter-process communication handlers.

Architecture Overview

The audio capture architecture in OpenWhispr follows a clear separation of concerns between the renderer and main processes. The renderer process handles all MediaRecorder interactions and UI state management, while the main process manages file system operations and transcription engine orchestration.

At the core of this system lies the AudioManager class defined in src/helpers/audioManager.js, which wraps the native MediaRecorder interface. React components interact with this manager through the useAudioRecording hook located in src/hooks/useAudioRecording.js, ensuring that UI logic remains decoupled from low-level audio APIs.

Step-by-Step Audio Capture Workflow

Triggering Recording via useAudioRecording Hook

When a user initiates dictation—either through a global hotkey or push-to-talk interface—the useAudioRecording hook invokes audioManagerRef.current.startRecording(). For streaming-enabled providers, the hook alternatively calls startStreamingRecording() to enable real-time partial transcription.

This action originates from React event handlers and propagates through the hook to the AudioManager instance, which maintains the active MediaRecorder reference throughout the recording session.

MediaRecorder Instantiation in AudioManager

Inside src/helpers/audioManager.js, the startRecording() method acquires a MediaStream through navigator.mediaDevices.getUserMedia({ audio: true }). The AudioManager then instantiates a new MediaRecorder instance with a MIME type of audio/webm to ensure broad browser compatibility.

// src/helpers/audioManager.js
async startRecording() {
  this.mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
  this.mediaRecorder = new MediaRecorder(this.mediaStream, { 
    mimeType: "audio/webm" 
  });
  this.chunks = [];
  
  // Event handlers attached here...
}

Blob Collection and Audio Level Monitoring

As the MediaRecorder captures audio, it emits dataavailable events containing Blob data. The AudioManager pushes each Blob into an internal chunks array for later concatenation. Simultaneously, the useAudioRecording hook can access real-time audio levels through audioManagerRef.current.getRecordingAudioLevel(), enabling the UI to render live audio visualization meters during dictation.

this.mediaRecorder.ondataavailable = (e) => {
  if (e.data.size) this.chunks.push(e.data);
};

Stopping and Concatenating Audio Chunks

When the user releases the hotkey or clicks the stop button, AudioManager.stopRecording() halts the MediaRecorder. The method returns a Promise that resolves once the onstop event fires, at which point the accumulated Blob chunks are merged into a single Blob and converted to an ArrayBuffer for efficient IPC transfer.

async stopRecording() {
  return new Promise((resolve) => {
    this.mediaRecorder.onstop = async () => {
      const blob = new Blob(this.chunks, { type: "audio/webm" });
      const buffer = await blob.arrayBuffer();
      // Transfer to main process via IPC
      window.api.sendAudioBuffer(buffer);
      resolve();
    };
    this.mediaRecorder.stop();
  });
}

IPC Transfer to Main Process

The resulting ArrayBuffer crosses the Electron context bridge via the 'save-audio' IPC channel handled in src/main/ipcHandlers.js. The main process receives the raw audio buffer and writes it to a temporary file in the user's data directory, creating a persistent audio source for transcription engines while keeping the renderer process memory-unburdened.

Transcription and Cleanup

The temporary audio file is routed to the selected speech-to-text engine—whether local whisper.cpp (via src/helpers/whisper.js), NVIDIA Parakeet (via src/helpers/parakeet.js), or a cloud provider. After transcription completes, the main process deletes the temporary file to maintain privacy. The text result travels back through IPC to the renderer, where useAudioRecording updates the UI and optionally pastes the transcription into the active application window.

React Integration Example

Components consume the audio capture functionality through the useAudioRecording hook, which abstracts the MediaRecorder lifecycle into simple start/stop controls:

// Hook that starts a dictation session
import useAudioRecording from "@/hooks/useAudioRecording";

function DictationButton() {
  const {
    start,
    stop,
    isRecording,
    streamingPartialText,
    streamingFinalText,
  } = useAudioRecording();

  const handleKeyDown = () => start({ voiceAgentRequested: false });
  const handleKeyUp = () => stop();

  return (
    <button
      onMouseDown={handleKeyDown}
      onMouseUp={handleKeyUp}
      disabled={isRecording}
    >
      {isRecording ? "Listening…" : "Start Dictation"}
    </button>
  );
}

Streaming vs. Batch Recording Modes

OpenWhispr supports two distinct recording paradigms determined by AudioManager.shouldUseStreaming(). Streaming mode captures and transmits audio chunks incrementally to compatible engines like NVIDIA Parakeet (sherpa-onnx), enabling low-latency partial transcriptions. Batch mode waits for the entire recording to complete before sending the consolidated audio buffer, which is required for engines like whisper.cpp that process complete files rather than streams.

Privacy-First Audio Handling

The entire pipeline operates with privacy as the default: all audio processing occurs locally unless the user explicitly configures a cloud provider. The MediaRecorder implementation ensures that raw audio buffers never persist in the renderer process longer than necessary, and temporary files are immediately purged after transcription. This architecture prevents data leakage while maintaining the responsiveness required for dictation workflows.

Summary

  • OpenWhispr uses the MediaRecorder API in src/helpers/audioManager.js to capture audio streams with the audio/webm MIME type.
  • The useAudioRecording hook in src/hooks/useAudioRecording.js provides React components with start/stop controls and real-time audio level monitoring via getRecordingAudioLevel().
  • Audio data transfers from renderer to main process via the 'save-audio' IPC channel as ArrayBuffers, then writes to temporary files for transcription.
  • The system supports both streaming transcription (NVIDIA Parakeet via src/helpers/parakeet.js) and batch processing (whisper.cpp via src/helpers/whisper.js) through shouldUseStreaming() logic.
  • Privacy safeguards include immediate deletion of temporary audio files and on-device processing by default.

Frequently Asked Questions

How does OpenWhispr handle audio format compatibility across different platforms?

OpenWhispr standardizes on the audio/webm MIME type when instantiating the MediaRecorder in src/helpers/audioManager.js. This format provides broad compatibility across Electron's Chromium renderer while maintaining efficient compression. The resulting WebM Blobs are converted to ArrayBuffers for IPC transfer, ensuring the main process receives platform-agnostic raw audio data regardless of the operating system's native audio APIs.

Can OpenWhispr stream audio for real-time transcription?

Yes, when AudioManager.shouldUseStreaming() returns true—typically when using NVIDIA Parakeet (sherpa-onnx)—the system activates startStreamingRecording() instead of the standard batch recording. This mode processes audio chunks incrementally through src/helpers/parakeet.js, delivering partial transcription results with minimal latency. However, engines like whisper.cpp require complete audio files, forcing the system to use batch mode with temporary file storage.

Where does OpenWhispr store audio data during recording?

Audio data exists in three stages: first as discrete Blobs in the renderer's chunks array within AudioManager, then as a consolidated ArrayBuffer transferred via IPC to the main process, and finally as a temporary file on disk in the user's data directory. The temporary file is created exclusively for transcription engine consumption and is deleted immediately after processing completes, ensuring no persistent audio artifacts remain on the filesystem.

Is microphone access in OpenWhispr secure?

OpenWhispr requests microphone access exclusively through the standard navigator.mediaDevices.getUserMedia() browser API, subject to Electron's permissionDialog handler. The application implements defense-in-depth by keeping raw audio in the renderer process only during active recording, transferring it immediately to the main process, and never transmitting audio to external services unless the user explicitly selects a cloud transcription provider. Local engines like whisper.cpp process audio entirely on-device.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →