How OpenWhispr Uses FFmpeg for Audio Processing: Complete Technical Guide

OpenWhispr delegates all low-level audio transformations to FFmpeg through a centralized utility module at src/helpers/ffmpegUtils.js, handling binary resolution, WAV conversion, segmentation, and merging to normalize recordings for Whisper ASR processing.

OpenWhispr is an open-source dictation and meeting capture application that relies on FFmpeg for critical audio processing tasks. Instead of embedding heavy audio libraries directly into the Electron main process, the application delegates format normalization, chunking, and reconstruction to a dedicated helper layer. This architecture ensures consistent, cross-platform audio handling while abstracting FFmpeg's complexity from the transcription and meeting detection logic.

Locating the FFmpeg Binary

OpenWhispr ships with cross-platform FFmpeg binaries via the ffmpeg-static package. The getFFmpegPath() function in src/helpers/ffmpegUtils.js implements a cascading resolution strategy to locate the executable.

First, the function checks for the bundled binary, including the unpacked app.asar.unpacked path used in production Electron builds. If the bundled binary cannot be executed, it falls back to common system locations such as /usr/bin/ffmpeg on Linux or C:\ffmpeg\bin\ffmpeg.exe on Windows. Finally, it scans the directories listed in the PATH environment variable. The resolved path is cached to avoid redundant filesystem operations on subsequent calls.

Converting Audio to Standard WAV Format

The transcription engine expects 16-bit PCM WAV files at 16 kHz mono for optimal Whisper ASR performance. The convertToWav(inputPath, outputPath, {sampleRate, channels}) function spawns an FFmpeg process with precision-tuned arguments:

ffmpeg -i <input> -ar 16000 -ac 1 -c:a pcm_s16le -y <output>

This function returns a Promise that resolves upon successful conversion or rejects with detailed error information parsed from FFmpeg's stderr output. OpenWhispr invokes this utility whenever a user-recorded file—such as a WebM blob from MediaRecorder or an uploaded MP3—requires normalization before ASR processing begins.

// Normalise a recorded WebM file to the WAV format required by Whisper
const { convertToWav } = require("./src/helpers/ffmpegUtils");
await convertToWav("/tmp/recording.webm", "/tmp/recording.wav");

Splitting Long Recordings into Manageable Chunks

For large audio files typical of meeting recordings, OpenWhispr implements fixed-duration segmentation to prevent memory exhaustion and enable parallel processing. The splitAudioFile(inputPath, outputDir, {segmentDuration, audioBitrate}) function defaults to 10-minute segments (600 seconds) and executes an FFmpeg segment command:

ffmpeg -i <input> -f segment -segment_time 600 -c:a libmp3lame -b:a 128k -ar 16000 -ac 1 -y <outputDir>/chunk-%03d.mp3

The function streams FFmpeg's stderr to capture the source duration metadata, supports clean abortion on cancellation requests, and returns an object containing the generated chunk file paths alongside the total duration in seconds. This segmentation occurs transparently when meeting recordings exceed internal processing limits.

// Split a long meeting recording into 10-minute MP3 chunks
const { splitAudioFile } = require("./src/helpers/ffmpegUtils");
const { chunkPaths, durationSeconds } = await splitAudioFile(
  "/tmp/meeting.webm",
  "/tmp/chunks",
  { segmentDuration: 600, audioBitrate: "128k" }
);

Merging Audio Segments Back Together

When reconstructing continuous audio streams—such as after failed segment-wise transcription retries—the mergeAudioSegments(segments) function constructs a temporary FFmpeg filter graph. The utility builds a complex filter string that resamples each input to 16 kHz, formats channels to mono, and concatenates streams:

[0:a]aresample=16000,aformat=sample_fmts=fltp:channel_layouts=mono[s0];
[1:a]aresample=16000,aformat=sample_fmts=fltp:channel_layouts=mono[s1];
...
[s0][s1]...concat=n=<N>:v=0:a=1[out]

FFmpeg then encodes the concatenated output using the Opus codec (-c:a libopus -b:a 64k). The merged audio is returned as a Buffer suitable for immediate playback or storage, eliminating intermediate temporary files.

// Re-assemble processed audio segments back into a single Opus file
const { mergeAudioSegments } = require("./src/helpers/ffmpegUtils");
const mergedBuffer = await mergeAudioSegments([
  { buffer: fs.readFileSync("/tmp/chunks/chunk-001.mp3"), mimeType: "audio/mp3" },
  { buffer: fs.readFileSync("/tmp/chunks/chunk-002.mp3"), mimeType: "audio/mp3" },
]);
fs.writeFileSync("/tmp/meeting-merged.opus", mergedBuffer);

Supporting Helper Utilities for Audio Analysis

Beyond format conversion, src/helpers/ffmpegUtils.js exports specialized utilities for audio validation and analysis:

  • isWavFormat / parseWavFormat – Quick validation and header extraction to verify file integrity before processing.
  • wavToFloat32Samples – Converts PCM samples to 32-bit float arrays for signal processing operations.
  • computeFloat32RMS – Calculates root-mean-square amplitude values, consumed by the meeting microphone gate in src/helpers/meetingMicGate.js to determine whether audio chunks contain speech.
  • parseFfmpegDuration – Parses the "Duration:" line from FFmpeg stderr to extract source length without requiring additional ffprobe calls.

Integration Points Across the Codebase

Audio Capture Flow

After MediaRecorder finishes capturing raw WebM data, the file passes immediately to convertToWav before entering the Whisper transcription queue. This occurs in the src/helpers/audioManager.js module, which bridges the browser's MediaRecorder API with the Node.js processing layer.

Meeting Recording Pipeline

When recordings exceed the internal 15-second segment limit, the application invokes splitAudioFile to create independent processing units. Each segment undergoes transcription separately, with mergeAudioSegments handling reconstruction if retry logic requires reassembly of select portions.

Diagnostic Tooling

The test suite at test/helpers/ffmpegUtils.test.js verifies conversion accuracy, splitting precision, and merging integrity across Windows, macOS, and Linux platforms, ensuring consistent FFmpeg behavior regardless of host environment.

All FFmpeg calls utilize Promise-based wrappers with comprehensive error handling, providing UI components with clear failure signals when binaries are missing, permissions are insufficient, or conversion parameters fail validation.

Summary

  • OpenWhispr centralizes FFmpeg operations in src/helpers/ffmpegUtils.js, ensuring consistent audio normalization across platforms.
  • Binary resolution cascades from bundled ffmpeg-static binaries through system paths to PATH environment variables via getFFmpegPath().
  • Format standardization converts arbitrary inputs to 16-bit PCM 16 kHz mono WAV using convertToWav() for Whisper compatibility.
  • Segmentation strategy splits large meeting recordings into 10-minute chunks with splitAudioFile(), while mergeAudioSegments() reconstructs continuous Opus streams using complex filter graphs.
  • Signal analysis helpers like computeFloat32RMS enable real-time microphone gating without external dependencies.
  • Robust error handling wraps all child process spawns, parsing FFmpeg stderr to provide actionable error messages to the user interface.

Frequently Asked Questions

How does OpenWhispr locate FFmpeg on different operating systems?

OpenWhispr uses the getFFmpegPath() function to check the bundled ffmpeg-static binary first (including Electron's app.asar.unpacked path), then falls back to platform-specific common locations like /usr/bin/ffmpeg or C:\ffmpeg\bin\ffmpeg.exe, and finally searches the system's PATH environment variable. The resolved path is cached for performance.

What audio format does OpenWhispr require for Whisper transcription?

OpenWhispr normalizes all audio to 16-bit PCM WAV files with a 16 kHz sample rate and mono channel configuration using the convertToWav() function. This specific format maximizes compatibility with OpenAI's Whisper ASR backend while minimizing preprocessing overhead.

How does OpenWhispr handle audio files longer than 10 minutes?

The splitAudioFile() function segments long recordings into 10-minute chunks (configurable via segmentDuration) using FFmpeg's segment muxer with MP3 encoding. This prevents memory issues during transcription and enables parallel processing of meeting recordings, with metadata capturing total duration for reconstruction.

Can OpenWhispr merge multiple audio segments back into a single file?

Yes, the mergeAudioSegments() function constructs an FFmpeg filter complex that resamples each segment to 16 kHz mono, concatenates them using the concat filter, and outputs a single Opus-encoded Buffer. This supports reconstructing meeting recordings after partial transcription retries or selective editing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →