# How OpenWhispr Integrates with the MediaRecorder API for Audio Capture

> Discover how OpenWhispr integrates with the MediaRecorder API to capture microphone audio. Learn about React hooks, AudioManager, and Electron IPC for seamless audio streaming and transcription.

- Repository: [OpenWhispr/openwhispr](https://github.com/OpenWhispr/openwhispr)
- Tags: how-to-guide
- Published: 2026-09-06

---

**OpenWhispr captures microphone audio using the browser's MediaRecorder API in the renderer process, coordinating between React hooks, the AudioManager helper, and Electron's IPC layer to stream audio chunks to the main process for local or cloud-based transcription.**

OpenWhispr is an open-source, privacy-first dictation application built with Electron and React. The application leverages the **MediaRecorder API** to capture raw audio streams directly in the browser renderer, ensuring sensitive voice data never leaves the device unless a cloud provider is explicitly selected. This integration creates a seamless pipeline from microphone input to speech-to-text transcription through a tightly coupled system of hooks, helpers, and inter-process communication handlers.

## Architecture Overview

The audio capture architecture in OpenWhispr follows a clear separation of concerns between the renderer and main processes. The **renderer process** handles all MediaRecorder interactions and UI state management, while the **main process** manages file system operations and transcription engine orchestration.

At the core of this system lies the `AudioManager` class defined in [`src/helpers/audioManager.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/audioManager.js), which wraps the native MediaRecorder interface. React components interact with this manager through the `useAudioRecording` hook located in [`src/hooks/useAudioRecording.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/hooks/useAudioRecording.js), ensuring that UI logic remains decoupled from low-level audio APIs.

## Step-by-Step Audio Capture Workflow

### Triggering Recording via useAudioRecording Hook

When a user initiates dictation—either through a global hotkey or push-to-talk interface—the `useAudioRecording` hook invokes `audioManagerRef.current.startRecording()`. For streaming-enabled providers, the hook alternatively calls `startStreamingRecording()` to enable real-time partial transcription.

This action originates from React event handlers and propagates through the hook to the AudioManager instance, which maintains the active MediaRecorder reference throughout the recording session.

### MediaRecorder Instantiation in AudioManager

Inside [`src/helpers/audioManager.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/audioManager.js), the `startRecording()` method acquires a MediaStream through `navigator.mediaDevices.getUserMedia({ audio: true })`. The AudioManager then instantiates a new `MediaRecorder` instance with a MIME type of `audio/webm` to ensure broad browser compatibility.

```javascript
// src/helpers/audioManager.js
async startRecording() {
  this.mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
  this.mediaRecorder = new MediaRecorder(this.mediaStream, { 
    mimeType: "audio/webm" 
  });
  this.chunks = [];
  
  // Event handlers attached here...
}

```

### Blob Collection and Audio Level Monitoring

As the MediaRecorder captures audio, it emits `dataavailable` events containing Blob data. The AudioManager pushes each Blob into an internal `chunks` array for later concatenation. Simultaneously, the `useAudioRecording` hook can access real-time audio levels through `audioManagerRef.current.getRecordingAudioLevel()`, enabling the UI to render live audio visualization meters during dictation.

```javascript
this.mediaRecorder.ondataavailable = (e) => {
  if (e.data.size) this.chunks.push(e.data);
};

```

### Stopping and Concatenating Audio Chunks

When the user releases the hotkey or clicks the stop button, `AudioManager.stopRecording()` halts the MediaRecorder. The method returns a Promise that resolves once the `onstop` event fires, at which point the accumulated Blob chunks are merged into a single Blob and converted to an ArrayBuffer for efficient IPC transfer.

```javascript
async stopRecording() {
  return new Promise((resolve) => {
    this.mediaRecorder.onstop = async () => {
      const blob = new Blob(this.chunks, { type: "audio/webm" });
      const buffer = await blob.arrayBuffer();
      // Transfer to main process via IPC
      window.api.sendAudioBuffer(buffer);
      resolve();
    };
    this.mediaRecorder.stop();
  });
}

```

### IPC Transfer to Main Process

The resulting ArrayBuffer crosses the Electron context bridge via the `'save-audio'` IPC channel handled in [`src/main/ipcHandlers.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/main/ipcHandlers.js). The main process receives the raw audio buffer and writes it to a temporary file in the user's data directory, creating a persistent audio source for transcription engines while keeping the renderer process memory-unburdened.

### Transcription and Cleanup

The temporary audio file is routed to the selected speech-to-text engine—whether local **whisper.cpp** (via [`src/helpers/whisper.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/whisper.js)), **NVIDIA Parakeet** (via [`src/helpers/parakeet.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/parakeet.js)), or a cloud provider. After transcription completes, the main process deletes the temporary file to maintain privacy. The text result travels back through IPC to the renderer, where `useAudioRecording` updates the UI and optionally pastes the transcription into the active application window.

## React Integration Example

Components consume the audio capture functionality through the `useAudioRecording` hook, which abstracts the MediaRecorder lifecycle into simple start/stop controls:

```tsx
// Hook that starts a dictation session
import useAudioRecording from "@/hooks/useAudioRecording";

function DictationButton() {
  const {
    start,
    stop,
    isRecording,
    streamingPartialText,
    streamingFinalText,
  } = useAudioRecording();

  const handleKeyDown = () => start({ voiceAgentRequested: false });
  const handleKeyUp = () => stop();

  return (
    <button
      onMouseDown={handleKeyDown}
      onMouseUp={handleKeyUp}
      disabled={isRecording}
    >
      {isRecording ? "Listening…" : "Start Dictation"}
    </button>
  );
}

```

## Streaming vs. Batch Recording Modes

OpenWhispr supports two distinct recording paradigms determined by `AudioManager.shouldUseStreaming()`. **Streaming mode** captures and transmits audio chunks incrementally to compatible engines like NVIDIA Parakeet (sherpa-onnx), enabling low-latency partial transcriptions. **Batch mode** waits for the entire recording to complete before sending the consolidated audio buffer, which is required for engines like whisper.cpp that process complete files rather than streams.

## Privacy-First Audio Handling

The entire pipeline operates with privacy as the default: all audio processing occurs locally unless the user explicitly configures a cloud provider. The MediaRecorder implementation ensures that raw audio buffers never persist in the renderer process longer than necessary, and temporary files are immediately purged after transcription. This architecture prevents data leakage while maintaining the responsiveness required for dictation workflows.

## Summary

- OpenWhispr uses the MediaRecorder API in [`src/helpers/audioManager.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/audioManager.js) to capture audio streams with the `audio/webm` MIME type.
- The `useAudioRecording` hook in [`src/hooks/useAudioRecording.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/hooks/useAudioRecording.js) provides React components with start/stop controls and real-time audio level monitoring via `getRecordingAudioLevel()`.
- Audio data transfers from renderer to main process via the `'save-audio'` IPC channel as ArrayBuffers, then writes to temporary files for transcription.
- The system supports both streaming transcription (NVIDIA Parakeet via [`src/helpers/parakeet.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/parakeet.js)) and batch processing (whisper.cpp via [`src/helpers/whisper.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/whisper.js)) through `shouldUseStreaming()` logic.
- Privacy safeguards include immediate deletion of temporary audio files and on-device processing by default.

## Frequently Asked Questions

### How does OpenWhispr handle audio format compatibility across different platforms?

OpenWhispr standardizes on the `audio/webm` MIME type when instantiating the MediaRecorder in [`src/helpers/audioManager.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/audioManager.js). This format provides broad compatibility across Electron's Chromium renderer while maintaining efficient compression. The resulting WebM Blobs are converted to ArrayBuffers for IPC transfer, ensuring the main process receives platform-agnostic raw audio data regardless of the operating system's native audio APIs.

### Can OpenWhispr stream audio for real-time transcription?

Yes, when `AudioManager.shouldUseStreaming()` returns true—typically when using NVIDIA Parakeet (sherpa-onnx)—the system activates `startStreamingRecording()` instead of the standard batch recording. This mode processes audio chunks incrementally through [`src/helpers/parakeet.js`](https://github.com/OpenWhispr/openwhispr/blob/main/src/helpers/parakeet.js), delivering partial transcription results with minimal latency. However, engines like whisper.cpp require complete audio files, forcing the system to use batch mode with temporary file storage.

### Where does OpenWhispr store audio data during recording?

Audio data exists in three stages: first as discrete Blobs in the renderer's `chunks` array within AudioManager, then as a consolidated ArrayBuffer transferred via IPC to the main process, and finally as a temporary file on disk in the user's data directory. The temporary file is created exclusively for transcription engine consumption and is deleted immediately after processing completes, ensuring no persistent audio artifacts remain on the filesystem.

### Is microphone access in OpenWhispr secure?

OpenWhispr requests microphone access exclusively through the standard `navigator.mediaDevices.getUserMedia()` browser API, subject to Electron's permissionDialog handler. The application implements defense-in-depth by keeping raw audio in the renderer process only during active recording, transferring it immediately to the main process, and never transmitting audio to external services unless the user explicitly selects a cloud transcription provider. Local engines like whisper.cpp process audio entirely on-device.