How OpenWhispr Transfers Audio Chunks from the Renderer to the Main Process
OpenWhispr transfers audio chunks from the renderer to the main process via Electron's IPC bridge, using real-time streaming for immediate chunk-by-chunk delivery or batch proxy mode for complete recording transfer, depending on the transcription provider's architecture. The AudioManager class orchestrates this hand-off through window.electronAPI methods exposed in the preload script.
OpenWhispr is an Electron-based application that enforces a strict separation between UI interactions and computational heavy lifting. While the renderer process handles microphone access and media recording using the Web Audio API, transcription engines—whether local Whisper cpp instances, cloud APIs, or ONNX-based models—execute in the main process. This design requires a robust pipeline to move binary audio data across the Electron process boundary without degrading application performance.
Recording and Buffering in the Renderer Process
The capture pipeline originates in src/helpers/audioManager.js. Here, the AudioManager class initializes a MediaRecorder instance (lines 54-55) that captures raw audio from the user's microphone. As the recorder emits dataavailable events, the manager pushes each resulting Blob into an internal accumulation array:
// Located in src/helpers/audioManager.js
this.audioChunks = [];
// During active recording...
mediaRecorder.ondataavailable = (event) => {
this.audioChunks.push(event.data);
};
This buffering continues until the user stops recording, at which point the manager constructs a payload object containing the collected chunks array (line 1703) alongside configuration metadata such as language codes and custom dictionary prompts.
Provider-Specific Transfer Architectures
OpenWhispr implements two distinct transfer strategies based on whether the selected provider supports real-time streaming or requires batch processing. These modes are defined in the provider mapping configuration spanning lines 352-404 of src/helpers/audioManager.js.
Real-Time Streaming Transfer
For streaming providers like Deepgram, AssemblyAI, OpenAI Realtime, and Corti, audio chunks transfer immediately upon capture rather than waiting for recording completion. The provider map exposes IPC helper functions that bridge directly to the main process:
// src/helpers/audioManager.js - Streaming provider configuration
const STREAMING_PROVIDERS = {
deepgram: {
warmup: opts => window.electronAPI.deepgramStreamingWarmup(opts),
start: opts => window.electronAPI.deepgramStreamingStart(opts),
send: buf => window.electronAPI.deepgramStreamingSend(buf), // Immediate IPC transfer
finalize: () => window.electronAPI.deepgramStreamingFinalize()
}
};
When AudioManager detects a streaming configuration, it invokes the provider's send method for each chunk as it arrives, passing the raw ArrayBuffer to window.electronAPI.deepgramStreamingSend(buf) (lines 355-357). This enables real-time partial transcription updates without accumulating large buffers in memory.
Batch Proxy Transfer
For proxy providers like Tinfoil, Mistral, Gemini, and XAI, the complete recording transfers only after capture ends. These providers expose an ipc function that returns a specific main-process handler:
// Proxy provider mapping in src/helpers/audioManager.js
const PROXY_TRANSCRIPTION_PROVIDERS = {
tinfoil: {
displayName: "Tinfoil",
ipc: () => window.electronAPI?.proxyTinfoilTranscription,
buildPayload: ({ audioBuffer, language, dictionaryPrompt }) => ({
audioBuffer,
language,
prompt: dictionaryPrompt || undefined
})
}
};
Upon stopping recording, the manager calls the provider's IPC function with the entire audioChunks array, transferring the complete session as a single payload (lines 398-404).
Main Process Reception and Transcription
The main process receives these transfers in src/main/ipcHandlers.js, which registers IPC channels for both streaming and proxy providers. For proxy providers, the handler receives the binary buffers, writes them to a temporary file on disk, and then invokes the appropriate transcription backend:
// Conceptual implementation in src/main/ipcHandlers.js
ipcMain.handle('proxy-tinfoil-transcription', async (event, payload) => {
const { chunks, language, prompt } = payload;
// Reconstruct audio from chunks array
const tempPath = await writeBuffersToTempFile(chunks);
// Execute transcription engine
return await executeTranscription(tempPath, { language, prompt });
});
Streaming handlers maintain persistent connections to external APIs or local engines, feeding individual chunks as they arrive without intermediate file system operations, enabling real-time partial result callbacks like onDeepgramPartialTranscript (lines 60-63).
Preload Bridge Security
The IPC bridge is securely exposed through src/main/preload.js using Electron's contextBridge API. This approach explicitly whitelists only the necessary methods, preventing unauthorized access to Node.js APIs from the renderer:
// src/main/preload.js
contextBridge.exposeInMainWorld('electronAPI', {
deepgramStreamingSend: (buffer) =>
ipcRenderer.invoke('deepgram-streaming-send', buffer),
proxyTinfoilTranscription: (payload) =>
ipcRenderer.invoke('proxy-tinfoil-transcription', payload),
// Additional provider handlers...
});
Complete Transfer Implementation
This implementation demonstrates the full workflow from capture to cross-process transfer:
import { AudioManager } from './helpers/audioManager.js';
const audioMgr = new AudioManager();
async function beginCapture() {
// Starts MediaRecorder and populates this.audioChunks
await audioMgr.startRecording();
}
async function finalizeAndTranscribe() {
// Stops recorder and finalizes payload { chunks: this.audioChunks, ... }
await audioMgr.stopRecording();
if (audioMgr.shouldUseStreaming()) {
// Chunks were already sent via streamingProvider.send() during recording
await audioMgr.streamingProvider.finalize();
} else {
// Batch transfer to main process
const provider = audioMgr.getProxyProvider();
const transcription = await provider.ipc().invoke('transcribe', {
chunks: audioMgr.audioChunks,
language: audioMgr.sttConfig.language,
dictionaryPrompt: audioMgr.customDictionary
});
return transcription;
}
}
Summary
- Dual-mode transfer: OpenWhispr supports streaming (chunk-by-chunk) for real-time APIs and proxy (complete batch) for engines requiring full audio context.
- Renderer capture: The
AudioManagerclass insrc/helpers/audioManager.jsusesMediaRecorderto buffer audio intothis.audioChunks, then initiates transfer viawindow.electronAPI. - IPC bridge:
src/main/preload.jssecurely exposes provider-specific channels likedeepgramStreamingSendandproxyTinfoilTranscription. - Main processing:
src/main/ipcHandlers.jsreceives binary data, writes proxy-provider chunks to temporary files, and manages transcription engine invocation. - Performance optimization: Streaming providers minimize latency by forwarding buffers directly to APIs, while proxy providers optimize accuracy by processing complete recordings.
Frequently Asked Questions
Why does OpenWhispr transfer audio to the main process instead of transcribing in the renderer?
Transcription engines—particularly local Whisper cpp or ONNX models—consume substantial CPU and memory resources that would block the renderer's UI thread. Running these in the main process prevents interface freezing and provides access to Node.js APIs for file system operations, native addons, and subprocess management.
What distinguishes streaming providers from proxy providers in OpenWhispr?
Streaming providers (Deepgram, AssemblyAI, OpenAI Realtime) receive audio chunks immediately as they are recorded, enabling real-time transcription with partial results returned via IPC callbacks. Proxy providers (Tinfoil, Mistral, Gemini) receive the entire audioChunks array only after recording stops, processing the complete session as a single unit. These distinctions are hardcoded in the provider maps at lines 352-388 and 398-404 of src/helpers/audioManager.js.
How does OpenWhispr handle large audio recordings during IPC transfer?
For proxy providers, the main process receives the chunks array as Blob objects and reconstructs the audio stream into a temporary file before transcription. This prevents memory exhaustion that could occur from holding large buffers in the renderer's heap. Streaming providers avoid this issue entirely by processing and discarding chunks incrementally without accumulation.
Can developers add custom transcription providers with different transfer methods?
Yes. Developers can extend src/helpers/audioManager.js by adding entries to either STREAMING_PROVIDERS (implementing warmup, start, send, and finalize methods) or PROXY_TRANSCRIPTION_PROVIDERS (implementing ipc and buildPayload methods). Corresponding handlers must be registered in src/main/ipcHandlers.js and exposed through src/main/preload.js via contextBridge to complete the IPC pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →