How OpenWhispr Integrates with the MediaRecorder API for Audio Capture
OpenWhispr captures microphone audio using the browser's MediaRecorder API in the renderer process, coordinating between React hooks, the AudioManager helper, and Electron's IPC layer to stream audio chunks to the main process for local or cloud-based transcription.
OpenWhispr is an open-source, privacy-first dictation application built with Electron and React. The application leverages the MediaRecorder API to capture raw audio streams directly in the browser renderer, ensuring sensitive voice data never leaves the device unless a cloud provider is explicitly selected. This integration creates a seamless pipeline from microphone input to speech-to-text transcription through a tightly coupled system of hooks, helpers, and inter-process communication handlers.
Architecture Overview
The audio capture architecture in OpenWhispr follows a clear separation of concerns between the renderer and main processes. The renderer process handles all MediaRecorder interactions and UI state management, while the main process manages file system operations and transcription engine orchestration.
At the core of this system lies the AudioManager class defined in src/helpers/audioManager.js, which wraps the native MediaRecorder interface. React components interact with this manager through the useAudioRecording hook located in src/hooks/useAudioRecording.js, ensuring that UI logic remains decoupled from low-level audio APIs.
Step-by-Step Audio Capture Workflow
Triggering Recording via useAudioRecording Hook
When a user initiates dictation—either through a global hotkey or push-to-talk interface—the useAudioRecording hook invokes audioManagerRef.current.startRecording(). For streaming-enabled providers, the hook alternatively calls startStreamingRecording() to enable real-time partial transcription.
This action originates from React event handlers and propagates through the hook to the AudioManager instance, which maintains the active MediaRecorder reference throughout the recording session.
MediaRecorder Instantiation in AudioManager
Inside src/helpers/audioManager.js, the startRecording() method acquires a MediaStream through navigator.mediaDevices.getUserMedia({ audio: true }). The AudioManager then instantiates a new MediaRecorder instance with a MIME type of audio/webm to ensure broad browser compatibility.
// src/helpers/audioManager.js
async startRecording() {
this.mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
this.mediaRecorder = new MediaRecorder(this.mediaStream, {
mimeType: "audio/webm"
});
this.chunks = [];
// Event handlers attached here...
}
Blob Collection and Audio Level Monitoring
As the MediaRecorder captures audio, it emits dataavailable events containing Blob data. The AudioManager pushes each Blob into an internal chunks array for later concatenation. Simultaneously, the useAudioRecording hook can access real-time audio levels through audioManagerRef.current.getRecordingAudioLevel(), enabling the UI to render live audio visualization meters during dictation.
this.mediaRecorder.ondataavailable = (e) => {
if (e.data.size) this.chunks.push(e.data);
};
Stopping and Concatenating Audio Chunks
When the user releases the hotkey or clicks the stop button, AudioManager.stopRecording() halts the MediaRecorder. The method returns a Promise that resolves once the onstop event fires, at which point the accumulated Blob chunks are merged into a single Blob and converted to an ArrayBuffer for efficient IPC transfer.
async stopRecording() {
return new Promise((resolve) => {
this.mediaRecorder.onstop = async () => {
const blob = new Blob(this.chunks, { type: "audio/webm" });
const buffer = await blob.arrayBuffer();
// Transfer to main process via IPC
window.api.sendAudioBuffer(buffer);
resolve();
};
this.mediaRecorder.stop();
});
}
IPC Transfer to Main Process
The resulting ArrayBuffer crosses the Electron context bridge via the 'save-audio' IPC channel handled in src/main/ipcHandlers.js. The main process receives the raw audio buffer and writes it to a temporary file in the user's data directory, creating a persistent audio source for transcription engines while keeping the renderer process memory-unburdened.
Transcription and Cleanup
The temporary audio file is routed to the selected speech-to-text engine—whether local whisper.cpp (via src/helpers/whisper.js), NVIDIA Parakeet (via src/helpers/parakeet.js), or a cloud provider. After transcription completes, the main process deletes the temporary file to maintain privacy. The text result travels back through IPC to the renderer, where useAudioRecording updates the UI and optionally pastes the transcription into the active application window.
React Integration Example
Components consume the audio capture functionality through the useAudioRecording hook, which abstracts the MediaRecorder lifecycle into simple start/stop controls:
// Hook that starts a dictation session
import useAudioRecording from "@/hooks/useAudioRecording";
function DictationButton() {
const {
start,
stop,
isRecording,
streamingPartialText,
streamingFinalText,
} = useAudioRecording();
const handleKeyDown = () => start({ voiceAgentRequested: false });
const handleKeyUp = () => stop();
return (
<button
onMouseDown={handleKeyDown}
onMouseUp={handleKeyUp}
disabled={isRecording}
>
{isRecording ? "Listening…" : "Start Dictation"}
</button>
);
}
Streaming vs. Batch Recording Modes
OpenWhispr supports two distinct recording paradigms determined by AudioManager.shouldUseStreaming(). Streaming mode captures and transmits audio chunks incrementally to compatible engines like NVIDIA Parakeet (sherpa-onnx), enabling low-latency partial transcriptions. Batch mode waits for the entire recording to complete before sending the consolidated audio buffer, which is required for engines like whisper.cpp that process complete files rather than streams.
Privacy-First Audio Handling
The entire pipeline operates with privacy as the default: all audio processing occurs locally unless the user explicitly configures a cloud provider. The MediaRecorder implementation ensures that raw audio buffers never persist in the renderer process longer than necessary, and temporary files are immediately purged after transcription. This architecture prevents data leakage while maintaining the responsiveness required for dictation workflows.
Summary
- OpenWhispr uses the MediaRecorder API in
src/helpers/audioManager.jsto capture audio streams with theaudio/webmMIME type. - The
useAudioRecordinghook insrc/hooks/useAudioRecording.jsprovides React components with start/stop controls and real-time audio level monitoring viagetRecordingAudioLevel(). - Audio data transfers from renderer to main process via the
'save-audio'IPC channel as ArrayBuffers, then writes to temporary files for transcription. - The system supports both streaming transcription (NVIDIA Parakeet via
src/helpers/parakeet.js) and batch processing (whisper.cpp viasrc/helpers/whisper.js) throughshouldUseStreaming()logic. - Privacy safeguards include immediate deletion of temporary audio files and on-device processing by default.
Frequently Asked Questions
How does OpenWhispr handle audio format compatibility across different platforms?
OpenWhispr standardizes on the audio/webm MIME type when instantiating the MediaRecorder in src/helpers/audioManager.js. This format provides broad compatibility across Electron's Chromium renderer while maintaining efficient compression. The resulting WebM Blobs are converted to ArrayBuffers for IPC transfer, ensuring the main process receives platform-agnostic raw audio data regardless of the operating system's native audio APIs.
Can OpenWhispr stream audio for real-time transcription?
Yes, when AudioManager.shouldUseStreaming() returns true—typically when using NVIDIA Parakeet (sherpa-onnx)—the system activates startStreamingRecording() instead of the standard batch recording. This mode processes audio chunks incrementally through src/helpers/parakeet.js, delivering partial transcription results with minimal latency. However, engines like whisper.cpp require complete audio files, forcing the system to use batch mode with temporary file storage.
Where does OpenWhispr store audio data during recording?
Audio data exists in three stages: first as discrete Blobs in the renderer's chunks array within AudioManager, then as a consolidated ArrayBuffer transferred via IPC to the main process, and finally as a temporary file on disk in the user's data directory. The temporary file is created exclusively for transcription engine consumption and is deleted immediately after processing completes, ensuring no persistent audio artifacts remain on the filesystem.
Is microphone access in OpenWhispr secure?
OpenWhispr requests microphone access exclusively through the standard navigator.mediaDevices.getUserMedia() browser API, subject to Electron's permissionDialog handler. The application implements defense-in-depth by keeping raw audio in the renderer process only during active recording, transferring it immediately to the main process, and never transmitting audio to external services unless the user explicitly selects a cloud transcription provider. Local engines like whisper.cpp process audio entirely on-device.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →