FluidVoice Limitations and Known Issues: A Complete Technical Guide
FluidVoice is a macOS-only dictation app with Apple Silicon hardware requirements, large model downloads, and mandatory system permissions that create specific constraints for users on older Macs or without cloud AI access.
FluidVoice delivers fast, private speech-to-text using a hybrid architecture of local CoreML models and an optional on-device AI layer called Fluid Intelligence. While this design prioritizes privacy and low latency, several architectural decisions in the altic-dev/FluidVoice codebase impose real limitations on platform support, hardware compatibility, and feature availability. This guide examines each constraint with direct reference to the source implementation.
macOS Platform Lock-In
FluidVoice cannot run on iOS, iPadOS, or Windows due to deep dependencies on macOS-specific frameworks.
- AppKit UI layer: The main interface in
Sources/Fluid/ContentView.swiftandSources/Fluid/fluidApp.swiftuses AppKit components unavailable on other platforms. - Accessibility typing API:
TypingService.swiftat line 260 injects text viaCGEventand the macOS Accessibility API—functionality with no equivalent on iOS or Windows. - Speech framework binding: The audio capture and speech recognition pipelines integrate with
NSSpeechRecognizerand Core Audio APIs exclusive to macOS.
Users requesting cross-platform support would require a complete UI rewrite and alternative text-injection mechanisms for each target platform.
Apple Silicon Hardware Requirement
Most bundled speech models are compiled for ARM64 and fail to execute on Intel Macs.
| Model | Architecture | Intel Compatibility |
|---|---|---|
| Nemotron Speech 3.5 | Apple Silicon only | ❌ Not supported |
| Parakeet Flash/TDT | Apple Silicon only | ❌ Not supported |
| Whisper (various sizes) | Multi-arch (ARM64 + x86_64) | ✅ Available fallback |
According to ASRService.swift at line 4969, the Parakeet Flash model implements a low-latency pipeline that assumes Apple Silicon-accelerated inference. Intel Mac users experience:
- Higher transcription latency
- Increased CPU utilization
- No access to the fastest local models
The codebase detects CPU architecture at runtime in ASRService.swift and filters available models accordingly.
Storage and Download Constraints
FluidVoice minimizes initial install size by downloading models on first use, but this creates significant storage demands:
// From ModelDownloader.swift, lines 412-414
// CoreML model binaries range from ~150MB to ~800MB per model
// Fluid Intelligence adds approximately 3.5GB
Download behavior:
- Models are cached in
~/Library/Caches/com.altic.FluidVoice/ - The
ModelDownloaderclass atSources/Fluid/Networking/ModelDownloader.swifthandles resume-able downloads with progress callbacks - No offline functionality without completed downloads
Users on metered connections or with limited storage must plan for multi-gigabyte transfers before full functionality is available.
Mandatory System Permissions
Two permissions are non-negotiable for core functionality:
- Microphone access: Required for audio capture
- Accessibility access: Required for text injection via
TypingService
Without accessibility permission, TypingService.swift line 260 cannot execute CGEventPost to type into other applications. The app detects this condition and surfaces it through:
// CommandModeView.swift displays this when prerequisites fail
if let issue = settings.commandModeReadinessIssue {
// Present issue #748 or permission-specific error
}
The commandModeReadinessIssue property aggregates permission states and known configuration problems for unified error presentation.
Cloud AI Dependency for Enhanced Features
Fluid Intelligence provides on-device AI post-processing, but when disabled, the app requires external API keys:
| Enhancement Type | Requirement | Data Handling |
|---|---|---|
| Fluid Intelligence (local) | ~3.5GB download, Apple Silicon | Fully private, no network |
| OpenAI/Groq fallback | API key stored in Keychain | Sent to provider |
From README lines 70-71: The app cannot perform AI-enhanced transcription without either local Fluid Intelligence or a valid cloud provider key. Raw transcription without enhancement remains available via local speech models.
Beta Channel Instability
An optional beta update stream exists at Settings → Automatic Updates → Beta Releases. Per README lines 71-73:
- Beta builds contain unfinished features
- Regressions are more frequent than stable releases
- New capabilities are validated on the main channel before beta promotion
Users requiring production stability should disable automatic beta enrollment.
Documented Bugs and Edge Cases
Two tracked issues affect specific usage patterns:
Issue #748: Window layout reset during dictation
- Triggered when
TypingServiceinjects text during certain app-focus transitions - Referenced in code comments at
TypingService.swiftline 260 - Workaround: Disable "Smart Formatting" in settings
Issue #445: URL parsing edge cases
- Malformed URLs in transcription output cause injection failures
- Test coverage exists in
LLMClientRequestBodyTests.swiftlines 204-210 - Affects Command Mode when dictating web addresses
Performance Characteristics on Older Hardware
Latency measurements vary significantly by hardware generation:
Parakeet Flash on M3 Max: ~40ms end-to-end
Parakeet Flash on M1: ~80ms end-to-end
Whisper Small on Intel i7: ~400-800ms end-to-end
Whisper Large on Intel i7: 2-4 seconds end-to-end
The ASRService streaming implementation in ASRService.swift optimizes for Apple Silicon memory bandwidth and neural engine availability. Intel Macs lack hardware acceleration for the primary model architectures.
Checking System Compatibility Programmatically
Developers integrating with or debugging FluidVoice can inspect runtime constraints:
import Fluid
// Check for ready-to-use Command Mode
let commandService = CommandModeService.shared
do {
try commandService.enterCommandMode()
} catch let error as LLMError {
// Handle .invalidRequest(issue) for missing permissions
print("Command Mode unavailable: \(error)")
}
// Verify model availability
do {
try ASRService.shared.startTranscription()
} catch {
// ModelNotFound – prompt download via Settings
print("Download required: \(error)")
}
The CommandModeService.shared singleton references settings.commandModeReadinessIssue to surface configuration problems before attempting transcription.
Summary
- Platform: macOS 15+ only—no iOS, iPadOS, or Windows support due to AppKit and Accessibility API dependencies
- Hardware: Apple Silicon strongly recommended; Intel Macs limited to slower Whisper models
- Storage: Plan for 3.5GB+ if using Fluid Intelligence; individual models require 150-800MB each
- Permissions: Microphone and Accessibility access are mandatory for typing functionality
- AI enhancement: Requires either local Fluid Intelligence download or external API key
- Stability: Beta channel available but with higher regression risk
- Known bugs: Issue #748 (layout reset) and #445 (URL parsing) affect specific workflows
Frequently Asked Questions
Can I run FluidVoice on my Intel Mac?
Intel Macs are partially supported. You can run Whisper-based models, but Nemotron Speech 3.5 and Parakeet Flash/TDT require Apple Silicon. The ASRService.swift architecture detection automatically filters incompatible models from the selection UI. Expect higher latency and CPU usage compared to Apple Silicon performance.
Why does FluidVoice require Accessibility permission?
The TypingService.swift implementation at line 260 uses CGEventPost to simulate keystrokes in other applications. This macOS Accessibility API allows the app to type transcription output anywhere you place your cursor. Without this permission, FluidVoice can record audio but cannot inject text, rendering it non-functional for dictation purposes.
How much storage space do I need for full functionality?
Minimum requirements vary by configuration: base app (~50MB), one speech model (150-800MB), and optionally Fluid Intelligence (~3.5GB). The ModelDownloader.swift implementation stores all files in user cache directories with automatic cleanup policies. First-time setup may require 4+ GB of available space for the complete feature set.
What happens if I don't enable Fluid Intelligence or add a cloud AI key?
Without either enhancement layer, FluidVoice performs raw speech-to-text transcription using local models only. You lose AI-powered features like automatic punctuation, formatting correction, and context-aware text refinement. The core dictation functionality remains fully operational through Nemotron, Parakeet, or Whisper models depending on your hardware.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →