How OpenSuperWhisper Handles Audio File Drag-and-Drop with Queue Processing
OpenSuperWhisper uses a SwiftUI FileDropOverlay modifier to capture audio file drops, extracts URLs via NSItemProvider, and feeds them into a sequential TranscriptionQueue that processes each file through Whisper one at a time.
OpenSuperWhisper is a macOS transcription app built in SwiftUI that streamlines audio-to-text conversion. The repository Starmel/OpenSuperWhisper implements a seamless audio file drag-and-drop with queue processing system that allows users to drop multiple files directly onto the main window, automatically queuing them for sequential transcription without manual intervention.
Detecting Audio File Drops in SwiftUI
The drag-and-drop interface centers on FileDropHandler.swift, which provides a view modifier that wraps the main content. While users drag files over the window, the modifier toggles isDragging to display a translucent overlay indicating that audio files can be dropped.
The drop detection uses SwiftUI's .onDrop modifier configured for UTType.audio. In OpenSuperWhisper/FileDropHandler.swift, the handler registers for audio content types and delegates to handleDrop(of:) when the user releases the mouse.
// OpenSuperWhisper/FileDropHandler.swift
.onDrop(of: [.audio], isTargeted: $handler.isDragging) { providers in
Task {
await handler.handleDrop(of: providers)
}
return true
}
Loading and Validating Dropped Files
When a drop occurs, the handleDrop(of:) method iterates through the supplied [NSItemProvider] array. For each provider conforming to UTType.audio, it asynchronously loads the underlying file URL using provider.loadItem(forTypeIdentifier:).
The code bridges the Objective-C completion handler into Swift's async/await pattern using withCheckedThrowingContinuation. If the URL loads successfully, the handler forwards it to the singleton TranscriptionQueue via await transcriptionQueue.addFileToQueue(url: url).
// OpenSuperWhisper/FileDropHandler.swift
func handleDrop(of providers: [NSItemProvider]) async {
for provider in providers where provider.hasItemConformingToTypeIdentifier(UTType.audio.identifier) {
do {
let url = try await withCheckedThrowingContinuation { cont in
provider.loadItem(forTypeIdentifier: UTType.audio.identifier) { item, err in
if let err = err { cont.resume(throwing: err); return }
cont.resume(returning: item as? URL)
}
}
guard let url = url else { continue }
await transcriptionQueue.addFileToQueue(url: url)
} catch {
print("Error loading dropped audio file: \(error)")
}
}
}
Queuing and Processing Pipeline
Creating Recording Objects
The TranscriptionQueue class, defined in OpenSuperWhisper/TranscriptionQueue.swift, manages the persistent state of pending transcriptions. When addFileToQueue(url:) receives a file, it constructs a Recording object by calculating the audio duration via AudioUtil.audioDuration(url:) and generating a unique UUID.
The recording stores metadata including the original source path (sourceFileURL), a timestamped filename, and an initial status of .pending. The method persists this record to RecordingStore before calling startProcessingQueue().
// OpenSuperWhisper/TranscriptionQueue.swift
func addFileToQueue(url: URL) async {
let duration = await AudioUtil.audioDuration(url: url)
let timestamp = Date()
let fileName = "\(Int(timestamp.timeIntervalSince1970)).wav"
let id = UUID()
let recording = Recording(
id: id,
timestamp: timestamp,
fileName: fileName,
transcription: "",
duration: duration,
status: .pending,
progress: 0.0,
sourceFileURL: url.path
)
try await recordingStore.addRecordingSync(recording)
startProcessingQueue()
}
Sequential Processing Loop
If the queue is not already running, startProcessingQueue() spawns a background Task that executes processQueue(). This method repeatedly queries RecordingStore for the next pending recording using getNextPendingRecording(), sets currentRecordingId, and invokes processRecording.
The processRecording method validates the source file exists, then runs the Whisper model through TranscriptionService. Progress updates are observed via setupProgressObserver and written back to the store, which drives the UI progress bars. When transcription completes or fails, the loop continues to the next item, ensuring sequential processing that preserves drop order.
// OpenSuperWhisper/TranscriptionQueue.swift
private func processQueue() async {
while let recording = recordingStore.getNextPendingRecording() {
currentRecordingId = recording.id
await processRecording(recording)
currentRecordingId = nil
}
}
Integration with the Main UI
To enable the workflow, the root ContentView in OpenSuperWhisper/ContentView.swift applies the .fileDropHandler() modifier. This single line attaches both the visual overlay and the drop handling logic, making the entire window surface receptive to audio files without additional buttons or import dialogs.
// OpenSuperWhisper/ContentView.swift
ContentView()
.fileDropHandler()
.sheet(isPresented: $isSettingsPresented) {
SettingsView()
}
Summary
- OpenSuperWhisper implements drag-and-drop through a custom
FileDropOverlaySwiftUI modifier inFileDropHandler.swift. - Dropped files are validated against
UTType.audioand loaded asynchronously viaNSItemProvider. - The
TranscriptionQueuesingleton createsRecordingobjects and persists them toRecordingStorebefore processing. - A background
TaskrunsprocessQueue()to transcribe files sequentially usingTranscriptionService. - Progress updates flow back to the UI through
setupProgressObserver, providing real-time feedback on transcription status.
Frequently Asked Questions
What file types does OpenSuperWhisper support for drag-and-drop?
The app specifically registers for UTType.audio identifiers in the NSItemProvider validation step within FileDropHandler.swift. This covers standard audio formats like WAV, MP3, and M4A that macOS recognizes as audio content, though the specific supported formats ultimately depend on the underlying Whisper model implementation in TranscriptionService.
How does the queue handle multiple files dropped at once?
When multiple files are dropped simultaneously, handleDrop(of:) iterates through all valid NSItemProvider objects in the array. Each URL is extracted and passed to addFileToQueue(url:) sequentially. The processQueue() method then processes them in the order they were added to the persistent store, maintaining the drop sequence.
Can users cancel a transcription after it enters the queue?
The queue tracks cancelled IDs and removes missing files automatically. While the specific cancellation API isn't detailed in the core flow, the TranscriptionQueue checks for cancelled states and validates file existence before processing each recording, allowing the loop to skip or terminate specific items gracefully.
Where does the app store intermediate recordings before processing?
Pending recordings are persisted immediately to RecordingStore, which maintains the queue state across app restarts. Each Recording object includes a sourceFileURL pointing to the original file location and a generated filename for internal reference, ensuring the queue can resume processing even if the app restarts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →