How OpenWhispr Collects Audio Data Chunks: A Two-Stage Pipeline for Low-Latency Dictation

OpenWhispr utilizes a dual-stage pipeline that combines pre-roll buffering with 250-millisecond chunked MediaRecorder slices to ensure reliable audio capture while minimizing cold-start latency.

This article examines the audio collection architecture implemented in the OpenWhispr repository, focusing on how the application balances raw microphone input, pre-roll safety buffers, and real-time streaming workflows. The strategy centers on src/helpers/audioManager.js, which orchestrates microphone constraints, blob chunking, and optional AudioWorklet processing to support both batch and streaming transcription backends.

The Core Strategy: Two-Stage Audio Capture

The collection system operates in two distinct phases to eliminate the "cold-start" bug where initial audio headers arrive without payload data. This approach ensures that even micro-utterances contain valid audio frames.

Stage 1: Pre-Roll Buffering with Prepared Microphone Capture

Before the user initiates dictation, OpenWhispr establishes a prepared microphone capture via preparedMicCapture.prepare. The system creates a MediaRecorder instance that records discreet chunks every RECORDING_TIMESLICE_MS = 250 milliseconds.

The ondataavailable handler stores each Blob in prepared.chunks, creating a rolling buffer that captures audio occurring immediately before the user presses the push-to-talk key. If the pre-roll exceeds PRE_ROLL_MAX_AGE_MS (3 seconds), the system discards stale buffers to prevent memory bloat and synchronization issues.

// Pre-roll handling – discard old pre-roll if it exceeds PRE_ROLL_MAX_AGE_MS (3s)
if (prepared && Date.now() - prepared.startedAt > PRE_ROLL_MAX_AGE_MS) {
  discardPreRoll(prepared);
}

Stage 2: Live Recording and Chunked MediaRecorder

When startRecording() fires, AudioManager either adopts the existing pre-roll stream or opens a fresh microphone stream. A new MediaRecorder consumes the MediaStream, continuing the 250-millisecond chunk interval and appending blobs to this.audioChunks.

This fixed interval guarantees that every utterance contains substantive audio data, fixing edge cases where header-only blobs previously caused transcription failures.

// Start a dictation recording (simplified)
await audioManager.startRecording(); // → prepares microphone, creates MediaRecorder
// … user speaks …
await audioManager.stopRecording(); // → concatenates all blobs, sends to transcription backend

Microphone Configuration for Raw Audio Input

In src/helpers/audioManager.js, the getAudioConstraints() method constructs a MediaStreamConstraints object that disables all browser-side audio processing. This ensures consistent raw input across Linux, macOS, and Windows environments.

The configuration explicitly sets:

  • echoCancellation: false
  • noiseSuppression: false
  • autoGainControl: false
  • channelCount: 2

The stereo channel requirement specifically addresses silence-detection problems on Linux systems running PipeWire, where mono streams occasionally fail to register activity.

// Constraints implemented in AudioManager.getAudioConstraints()
const constraints = {
  audio: {
    echoCancellation: false,
    noiseSuppression: false,
    autoGainControl: false,
    channelCount: 2
  }
};

Real-Time Preview via AudioWorklet

For live transcription preview, OpenWhispr attaches an AudioWorklet named pcm-streaming-processor to the active microphone stream. This worklet buffers PCM samples in 800-sample blocks, posting each complete buffer back to the main thread via port.postMessage.

The main thread forwards these blocks through window.electronAPI.sendDictationPreviewAudio, enabling parallel processing without interrupting the primary MediaRecorder stream.

// Create the preview worklet (run inside AudioManager)
const workletCode = `
class PCMStreamingProcessor extends AudioWorkletProcessor {
  constructor() {
    super();
    this._buffer = new Int16Array(800);
    this._offset = 0;
    this._stopped = false;
    this.port.onmessage = e => { if (e.data === "stop") this._stopped = true; };
  }
  process(inputs) {
    if (this._stopped) return false;
    const input = inputs[0]?.[0];
    if (!input) return true;
    for (let i = 0; i < input.length; i++) {
      const s = Math.max(-1, Math.min(1, input[i]));
      this._buffer[this._offset++] = s < 0 ? s * 0x8000 : s * 0x7fff;
      if (this._offset >= 800) {
        this.port.postMessage(this._buffer.buffer, [this._buffer.buffer]);
        this._buffer = new Int16Array(800);
        this._offset = 0;
      }
    }
    return true;
  }
}
registerProcessor("pcm-streaming-processor", PCMStreamingProcessor);
`;

Streaming-Commit Mode for NVIDIA Parakeet Models

When using local Whisper with NVIDIA Parakeet online models, the AudioWorklet drives a specialized streaming-commit mode. Instead of buffering for batch processing, the PCM stream flows directly to the streaming provider via the same sendDictationPreviewAudio channel.

This configuration triggers when useLocalWhisper && isNvidia && isOnlineParakeetModel(parakeetModel) evaluates to true, allowing transcription commits to occur simultaneously with speech rather than waiting for recording cessation.

Key Architectural Components

The audio collection strategy relies on several specialized helpers:

Summary

OpenWhispr's strategy for collecting audio data chunks centers on reliability and latency optimization:

  • Raw microphone input – Browser processing effects are disabled to prevent audio pipeline interference.
  • 250-millisecond chunking – Fixed-interval MediaRecorder slices ensure valid audio frames in every blob.
  • Pre-roll buffering – 3-second rolling buffers capture audio occurring before explicit record commands.
  • Parallel AudioWorklet – 800-sample PCM blocks feed real-time preview and streaming-commit models without blocking the main recording thread.

Frequently Asked Questions

Why does OpenWhispr disable echo cancellation and noise suppression?

OpenWhispr disables echoCancellation, noiseSuppression, and autoGainControl to ensure the transcription model receives unprocessed, consistent audio data. According to the source code in src/helpers/audioManager.js, browser-side processing can introduce artifacts that degrade Whisper model accuracy, particularly on Linux systems where audio pipelines vary significantly between distributions.

What is the purpose of the 250-millisecond recording timeslice?

The RECORDING_TIMESLICE_MS = 250 value ensures that MediaRecorder emits actionable blobs containing actual audio samples rather than just container headers. This interval fixes the "cold-start" bug where very short utterances previously resulted in empty blobs that caused transcription failures. The 250ms window provides a balance between latency and data density.

How does the pre-roll buffer handle microphone permission delays?

The preparedMicCapture system initializes the microphone stream before the user presses the record key, eliminating permission prompt delays during actual dictation. If the user delays speaking for more than PRE_ROLL_MAX_AGE_MS (3 seconds), OpenWhispr automatically discards the stale buffer and prepares a fresh capture to maintain synchronization between audio and transcription timestamps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →