Does FluidVoice Support Real-Time Speech Synthesis?
FluidVoice does not support real-time speech synthesis; it is built exclusively for real-time speech-to-text dictation and transcription.
FluidVoice is an open-source macOS voice dictation application focused entirely on converting spoken audio into written text. While it delivers robust real-time speech recognition through native macOS APIs, the codebase contains no text-to-speech (TTS) engine or speech synthesis functionality.
How FluidVoice Handles Speech Processing
The application architecture centers on three core services that manage live audio capture and streaming transcription. According to the altic-dev/FluidVoice source code, these components work together to deliver low-latency dictation without any synthesis capability.
ASRService.swift: The Central Recognition Pipeline
The ASRService class in Sources/Fluid/Services/ASRService.swift orchestrates the entire speech-to-text workflow. It manages audio buffering, session lifecycle, and delegates recognition tasks to platform-specific providers.
// Start a real-time transcription session
let asr = ASRService.shared
asr.startSession(language: .englishUS) { result in
switch result {
case .partial(let text):
// Update UI with interim transcription
print("Partial: \(text)")
case .final(let text):
// Finalized sentence
print("Final: \(text)")
case .error(let err):
print("Error: \(err)")
}
}
This callback-based API returns partial and final transcription results—never synthesized audio.
AppleSpeechProvider.swift: macOS Recognition Wrapper
Sources/Fluid/Services/AppleSpeechProvider.swift wraps macOS's built-in SFSpeechRecognizer for standard recognition tasks. As implemented in altic-dev/FluidVoice, this provider only handles incoming audio streams:
// AppleSpeechProvider – low-level call to macOS recognizer
let provider = AppleSpeechProvider()
provider.recognizeAudioStream(url: audioURL) { transcript in
// transcript contains the spoken words – no synthesis involved
print(transcript)
}
AppleSpeechAnalyzerProvider.swift: Streaming for Low Latency
For advanced use cases requiring faster partial results, Sources/Fluid/Services/AppleSpeechAnalyzerProvider.swift implements SFSpeechAnalyzer—Apple's streaming recognition API introduced for continuous dictation scenarios. This remains transcription-only with no synthesis path.
Evidence Against Speech Synthesis Support
A comprehensive review of the FluidVoice Swift source tree reveals no TTS implementation:
- No
AVSpeechSynthesizerimports or usage — Apple's standard synthesis framework is entirely absent - No "synthesize" method calls related to audio — the only occurrences reference UI progress bar updates
- No synthesis settings in
SettingsStore.swift— located atSources/Fluid/Persistence/SettingsStore.swift, this file only stores recognition model preferences (appleSpeech,appleSpeechAnalyzer) - No synthesis UI in
ContentView.swift— the main interface atSources/Fluid/ContentView.swiftdisplays transcription results and recording controls exclusively
What FluidVoice Actually Does
Rather than speech synthesis, FluidVoice provides these real-time speech-to-text features:
- Live streaming transcription with partial and final result callbacks
- Speaker diarization to identify different speakers in recorded audio
- Multiple recognition backends via
appleSpeechandappleSpeechAnalyzermodels - Continuous dictation sessions with automatic segmenting and punctuation
Summary
- FluidVoice is speech-to-text only — it converts audio to text, not text to audio
- The codebase contains zero synthesis infrastructure — no
AVSpeechSynthesizer, custom TTS engines, or audio generation pipelines - All speech-related services (
ASRService,AppleSpeechProvider,AppleSpeechAnalyzerProvider) handle recognition exclusively - Users seeking real-time speech synthesis need separate macOS TTS tools or alternative applications
Frequently Asked Questions
Can FluidVoice read text aloud?
No. FluidVoice has no text-to-speech capability. The application only captures microphone input and returns transcribed text. For macOS text-to-speech, use system features like Spoken Content in Accessibility settings or third-party TTS applications.
Does FluidVoice plan to add speech synthesis?
The current codebase shows no foundation for synthesis features. The architecture is purpose-built for recognition efficiency, with ASRService designed around transcription provider patterns that don't accommodate audio output generation.
What macOS speech APIs does FluidVoice use?
FluidVoice uses SFSpeechRecognizer (via AppleSpeechProvider) and SFSpeechAnalyzer (via AppleSpeechAnalyzerProvider)—both Apple frameworks for speech recognition only. These APIs handle audio input and return text transcripts; they do not perform synthesis.
Is there any speech generation in FluidVoice's audio pipeline?
No. The audio pipeline in ASRService.swift is unidirectional: microphone → buffer → recognition provider → text callback. No audio output path exists, and no synthesis classes are referenced anywhere in Sources/Fluid/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →