How the OpenSuperWhisper Transcription Queue Handles Multiple Audio Files
OpenSuperWhisper processes many recordings through a single-threaded, async FIFO queue that guarantees only one file is transcribed at a time while keeping the UI responsive.
When users drop multiple audio files into OpenSuperWhisper, the application must serialize access to the underlying transcription engine while maintaining a responsive interface. The solution is a robust queueing system implemented in TranscriptionQueue.swift that manages the lifecycle of each recording from intake to completion. Understanding how this transcription queue handles multiple audio files reveals the architecture behind its reliable FIFO processing.
Queue Entry Point and File Registration
When a new audio file enters the system, the addFileToQueue(url:) method creates a Recording object containing metadata like duration, temporary filename, and UUID. This object is immediately persisted to the SQLite-backed RecordingStore defined in OpenSuperWhisper/Models/RecordingStore.swift, ensuring durability across app restarts.
After persistence, startProcessingQueue() activates the pipeline. It employs a guard clause—guard !isProcessing else { return }—to prevent concurrent queue executions. The method first invokes cleanupMissingFiles() to remove stale entries, then initiates the async processQueue() task within a dedicated Task to isolate work from the main thread.
FIFO Sequential Processing
The processQueue() method in OpenSuperWhisper/TranscriptionQueue.swift implements strict first-in-first-out ordering. It queries recordingStore.getNextPendingRecording(), which filters for statuses like .pending, .converting, or .transcribing and orders results by timestamp ascending. This ensures recordings are pulled one-by-one in chronological order.
The queue maintains a while loop that continues processing as long as pending recordings exist. After each transcription completes, the loop clears currentRecordingId, removes the UUID from cancelledRecordingIds, and immediately proceeds to the next pending item. When the queue empties, isProcessing resets to false, allowing future calls to startProcessingQueue() to reactivate the pipeline.
Per-Recording Workflow and State Management
Each iteration calls processRecording(_:) to handle the specific lifecycle of a single audio file. This method performs several critical validations and transitions:
Cancellation Check. The system verifies if the recording ID exists in cancelledRecordingIds. If present, the task skips immediately.
Source Validation. The file path stored in sourceFileURL is verified to exist on disk. Missing files trigger an immediate status change to .failed.
Status Transitions. Valid recordings move through .converting (or .transcribing for regenerations) via atomic updates like recordingStore.updateRecordingProgressOnlySync or updateRecordingStatusOnly.
Transcription Execution. A Task assigned to currentTranscriptionTask invokes transcriptionService.transcribeAudio(url:settings:). The TranscriptionService serializes access to the Whisper or FluidAudio engine by awaiting any previous transcriptionTask before starting new work. Progress updates publish to TranscriptionService.progress, which TranscriptionQueue observes to update UI state.
Result Handling. Success moves temporary files to permanent storage—moving if from the app's temp folder or copying if user-supplied—then updates the Recording row with final text, progress = 1.0, and status .completed.
Error Handling. Any thrown error (except cancellation) marks the record as .failed with a descriptive message.
Concurrency Guarantees and Cancellation
The architecture enforces strict single-threaded execution despite the async Swift concurrency model. Only one currentTranscriptionTask exists at any moment, enforced by transcriptionTask guards within TranscriptionService. The queue runs inside processingTask, isolating async work from the main thread, while @Published properties like isProcessing and currentRecordingId always update on the main actor for UI consistency.
Cancellation operates granularly. Calling cancelRecording(id:) adds the UUID to cancelledRecordingIds. If the ID matches the active recording, the system aborts the engine via TranscriptionService.cancelTranscription() and cancels the internal task, allowing the queue to proceed to the next item without stopping the entire pipeline.
Implementation Examples
To add files programmatically from a UI drop handler:
// In your FileDropHandler.swift or similar
Task {
await TranscriptionQueue.shared.addFileToQueue(url: droppedFileURL)
}
To manually trigger queue processing (rarely needed as the queue auto-starts):
TranscriptionQueue.shared.startProcessingQueue()
To cancel a specific recording:
let idToCancel: UUID = // target recording ID
TranscriptionQueue.shared.cancelRecording(idToCancel)
Observing queue state in SwiftUI:
@ObservedObject private var queue = TranscriptionQueue.shared
var body: some View {
ProgressView(value: queue.isProcessing ? 0.5 : 0)
.opacity(queue.isProcessing ? 1 : 0)
}
Summary
- OpenSuperWhisper uses a single-threaded async queue in
TranscriptionQueue.swiftto process audio files sequentially. - The FIFO ordering is guaranteed by
recordingStore.getNextPendingRecording()filtering by timestamp. - Concurrency safety comes from
guard !isProcessingchecks and serialized access to the transcription engine viaTranscriptionService. - State persistence relies on SQLite-backed
RecordingStoreto maintain queue integrity across app sessions. - Fine-grained cancellation allows stopping individual recordings without halting the entire queue.
Frequently Asked Questions
Can OpenSuperWhisper transcribe multiple audio files simultaneously?
No. The transcription queue handles multiple audio files sequentially, not in parallel. The guard !isProcessing check in startProcessingQueue() and the serialized transcriptionTask in TranscriptionService ensure only one file is processed at a time, preventing resource contention on the Whisper engine.
What happens if I add files while transcription is already running?
The addFileToQueue(url:) method persists the new file to RecordingStore immediately. If the queue is already processing (isProcessing == true), the file waits in the database. The existing processQueue() loop will pick it up automatically when it reaches the end of the current FIFO sequence, or when the queue restarts after completion.
How does the queue maintain order when multiple files are dropped?
Files are processed in FIFO (First-In-First-Out) order. When processQueue() queries recordingStore.getNextPendingRecording(), it filters for pending statuses and sorts by timestamp ascending. This ensures the oldest unprocessed file is always selected next, regardless of how many files were added in bulk.
Where is the queue state stored if the app crashes?
Queue state persists in a SQLite database via RecordingStore. Each Recording object is saved to disk in OpenSuperWhisper/Models/RecordingStore.swift before transcription begins. On restart, cleanupMissingFiles() removes orphaned entries, and startProcessingQueue() can resume processing any remaining pending items.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →