How to Batch Process Audio Files with OpenSuperWhisper: Complete Guide
OpenSuperWhisper processes multiple audio files sequentially through a shared TranscriptionQueue that accepts files via drag-and-drop, the Open menu, or programmatic API calls, running transcriptions one-by-one in background tasks while keeping the UI responsive on the main actor.
OpenSuperWhisper provides a robust batch processing system for transcribing multiple audio files without manual intervention. The macOS application handles file queuing through a centralized TranscriptionQueue that ensures sequential processing, preventing resource contention while maintaining responsive user interface updates. Whether you import files via drag-and-drop or custom scripts, the underlying architecture guarantees that each recording completes before the next begins.
Understanding the Batch Processing Architecture
OpenSuperWhisper implements a single-threaded transcription queue that serializes all audio processing operations. When you drop one or more files onto the app window, FileDropHandler.handleDrop(of:) iterates over the NSItemProvider instances, extracts the URL of each audio file, and calls await TranscriptionQueue.shared.addFileToQueue(url:).
The queue runs on the main actor to ensure UI state consistency, while the heavy transcription work executes in Task.detached(priority:) background tasks. This design guarantees that no two recordings are transcribed simultaneously, avoiding resource contention and keeping the interface responsive.
Key source files in this flow:
OpenSuperWhisper/FileDropHandler.swift– Handles drag-and-drop extraction and URL forwardingOpenSuperWhisper/TranscriptionQueue.swift– Central queue logic and sequential processing driverOpenSuperWhisper/TranscriptionService.swift– Wraps Whisper/FluidAudio engines and publishes progress
Method 1: Drag-and-Drop Batch Processing
The simplest way to batch process audio files is using the built-in drag-and-drop interface. When you attach the file drop handler to your SwiftUI view, the app automatically manages multiple files.
// In any SwiftUI view, attach the file-drop handler:
struct ContentView: View {
var body: some View {
VStack {
// …your UI…
}
.fileDropHandler() // ← enables drag-and-drop
}
}
When the user drops several audio files, FileDropHandler.handleDrop(of:) automatically enqueues each file. The UI displays a translucent overlay with the message "Multiple files will be queued" (lines 68-71 of OpenSuperWhisper/FileDropHandler.swift). Each file enters the queue sequentially without overlapping processing.
Method 2: Programmatic Batch Enqueuing
For automation or custom workflows, you can programmatically add batches of URLs to the queue using the shared instance:
import OpenSuperWhisper
func enqueueBatch(_ urls: [URL]) async {
let queue = TranscriptionQueue.shared
for url in urls {
await queue.addFileToQueue(url: url)
}
}
Calling enqueueBatch([url1, url2, …]) places every file into the same queue, and TranscriptionQueue.startProcessingQueue() automatically begins processing if it isn't already running. The method addFileToQueue(url:) performs several operations for each file:
- Determines audio duration using
AudioUtil.audioDuration - Creates a
Recordingmodel with a temporary filename - Stores the model in
RecordingStore - Triggers
startProcessingQueue()if idle
How the Queue Processes Files
Once files are enqueued, TranscriptionQueue.processQueue() repeatedly queries RecordingStore.getNextPendingRecording() for the next pending item. For each recording, processRecording(_:) invokes TranscriptionService.transcribeAudio(url:settings:) to execute the actual Whisper inference.
The flow executes in this order:
startProcessingQueue()spawns a backgroundTaskprocessQueue()pulls recordings one-by-oneTranscriptionServiceruns the Whisper engine (or FluidAudio engine)- Progress updates feed back to the UI via
TranscriptionService.$progress - On completion, temporary files move to the permanent recordings directory
- When the queue empties,
isProcessingclears, allowing new batches to start
Monitoring Transcription Progress
Track batch processing status through observable properties on the shared queue:
struct ProgressView: View {
@ObservedObject private var queue = TranscriptionQueue.shared
var body: some View {
if queue.isProcessing {
Text("Transcribing…")
.progressViewStyle(.linear)
.opacity(0.8)
}
}
}
The isProcessing flag and per-recording progress updates originate from setupProgressObserver() within TranscriptionQueue. Because the queue operates on the main actor, these state changes reflect immediately in your SwiftUI views without thread-safety concerns.
Cancelling Active or Pending Transcriptions
To cancel a specific recording before or during processing, call the cancellation method with the recording's UUID:
let someRecordingId: UUID = …
TranscriptionQueue.shared.cancelRecording(someRecordingId)
The cancellation logic removes the ID from cancelledRecordingIds and aborts the running Task if it matches the current recording. This allows you to stop problematic files without clearing the entire batch queue.
Summary
- Use
TranscriptionQueue.sharedto enqueue multiple files viaaddFileToQueue(url:)from any context—drag-and-drop handlers, menu actions, or background scripts. - Files process sequentially through
processQueue()usingTask.detached(priority:)background tasks, ensuring only one transcription runs at a time. - Monitor progress through the
isProcessingboolean andTranscriptionServicepublishers that update the main actor. - Cancel specific jobs using
cancelRecording(_:)with the recording UUID to remove items fromcancelledRecordingIdswithout affecting the rest of the batch. - Source files implementing this behavior include
OpenSuperWhisper/FileDropHandler.swift,OpenSuperWhisper/TranscriptionQueue.swift, andOpenSuperWhisper/TranscriptionService.swift.
Frequently Asked Questions
How does OpenSuperWhisper handle multiple files simultaneously?
OpenSuperWhisper does not process files simultaneously. It uses a single-threaded queue (TranscriptionQueue) that processes files one-by-one through processRecording(_:). This sequential approach prevents resource contention and ensures the Whisper engine or FluidAudio engine handles only one audio file at a time.
Can I add files to the queue while transcription is already running?
Yes, the shared queue accepts new files at any time via addFileToQueue(url:). If isProcessing is true, the new files wait in RecordingStore until the current transcription completes. Once the active job finishes, processQueue() automatically picks up the next pending recording.
Where does OpenSuperWhisper store temporary files during batch processing?
When addFileToQueue(url:) creates a Recording model, it generates a temporary filename. The file remains in temporary storage during transcription, then moves (or copies, depending on origin) into the permanent recordings directory upon successful completion, as implemented in OpenSuperWhisper/TranscriptionQueue.swift.
How can I check which file is currently being transcribed?
Monitor TranscriptionQueue.shared.isProcessing to determine if transcription is active. For detailed progress, observe TranscriptionService.$progress, which publishes updates during transcribeAudio(url:settings:). The queue updates these properties on the main actor, ensuring immediate UI synchronization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →