Read Frog Text-to-Speech Implementation: Edge TTS Voice Selection and Speed Control
Read Frog does not integrate OpenAI's Text-to-Speech API; instead, it implements TTS through Microsoft Edge TTS with automatic language-based voice selection and granular control over speech rate, pitch, and volume.
While many modern reading applications integrate OpenAI's Text-to-Speech service for neural voice generation, the open-source Read Frog browser extension takes a different architectural approach. According to the source code in mengxi-ream/read-frog, the project deliberately avoids OpenAI TTS in favor of Microsoft's Edge TTS API, which provides robust multilingual support without requiring separate API subscriptions. This implementation offers sophisticated voice selection logic based on automatic language detection and fine-grained audio controls.
Architecture Overview
The TTS pipeline in Read Frog operates through a six-stage process orchestrated by the useTextToSpeech React hook in src/hooks/use-text-to-speech.tsx:
- Trigger – UI components invoke the
playmethod with text content andTTSConfig. - Voice Resolution –
resolveVoiceForTextdetects the input language viadetectLanguageand selects a voice fromttsConfig.languageVoices, falling back tottsConfig.defaultVoiceif no language-specific mapping exists. - Chunking –
splitTextByUtf8Bytesdivides text into chunks ≤ 2KB UTF-8 bytes to respect API limits. - Synthesis Request – Each chunk triggers
sendMessage("edgeTtsSynthesize", …)with payload parameters includingvoice,rate,pitch, andvolume. - Audio Playback – The background script
src/entrypoints/background/tts-playback.tscreates an off-screen document to bypass Content Security Policy restrictions and forwards base64 audio viattsOffscreenPlay. - Caching – React Query caches generated audio chunks under
["tts-audio", …]keys to prevent redundant synthesis.
Voice Selection Logic
Read Frog implements data-driven voice selection through the TTSConfig type defined in src/types/config/tts.ts. Rather than relying on OpenAI's limited voice roster, the system maps ISO-639-3 language codes to specific Edge TTS voices:
- Language Detection – The
detectLanguageutility analyzes input text to determine its language code. - Voice Mapping – The system queries
ttsConfig.languageVoicesusing the detected code. If a mapping exists (e.g., "eng" → "en-US-AriaNeural"), that voice is selected. - Fallback Strategy – When no language-specific voice is configured or detection fails, the system defaults to
ttsConfig.defaultVoice.
This approach allows users to configure distinct voices for different languages, enabling multilingual reading sessions with appropriate native pronunciation.
Speech Speed and Audio Controls
While OpenAI TTS offers limited speed control (typically 0.25x to 4.0x multipliers), Read Frog provides granular adjustment of three audio parameters through the TTSConfig interface:
- Rate (speech speed): Range –100 to +100, converted to percentage strings (e.g.,
"+10%"or"-20%") - Pitch: Range –100 to +100, converted to Hertz strings (e.g.,
"+5Hz") - Volume: Range –100 to +100, converted to percentage strings
The toSignedValue utility function in src/hooks/use-text-to-speech.tsx handles the transformation:
function toSignedValue(value: number, unit: "%" | "Hz"): string {
return `${value >= 0 ? "+" : ""}${value}${unit}`
}
These values are passed to the Edge TTS API through the sendMessage("edgeTtsSynthesize", …) payload, allowing real-time adjustment of speech characteristics without modifying the underlying audio files.
Code Implementation Examples
React Hook Usage
import { useTextToSpeech } from "@/hooks/use-text-to-speech"
import { useConfig } from "@/hooks/use-config"
function SpeakButton({ text }: { text: string }) {
const { config } = useConfig()
const { play, stop, isPlaying, currentChunk, totalChunks } = useTextToSpeech()
const handleClick = async () => {
await play(text, config.tts) // config.tts implements TTSConfig
}
return (
<button onClick={handleClick} disabled={isPlaying}>
{isPlaying ? `Playing ${currentChunk}/${totalChunks}` : "Speak"}
</button>
)
}
Source: src/hooks/use-text-to-speech.tsx
Voice Resolution Implementation
async function resolveVoiceForText(text: string, ttsConfig: TTSConfig): Promise<string> {
const detectedLanguage = await detectLanguage(text, {
minLength: 0,
enableLLM: ttsConfig.detectLanguageMode === "llm",
})
if (detectedLanguage && detectedLanguage in ttsConfig.languageVoices) {
return ttsConfig.languageVoices[detectedLanguage as keyof typeof ttsConfig.languageVoices] ?? ttsConfig.defaultVoice
}
return ttsConfig.defaultVoice
}
Source: src/hooks/use-text-to-speech.tsx
Edge TTS Synthesis Request
const response = await sendMessage("edgeTtsSynthesize", {
text: chunk,
voice,
rate: toSignedValue(ttsConfig.rate, "%"),
pitch: toSignedValue(ttsConfig.pitch, "Hz"),
volume: toSignedValue(ttsConfig.volume, "%"),
})
Source: src/hooks/use-text-to-speech.tsx
Why Read Frog Uses Edge TTS Instead of OpenAI
The Read Frog codebase contains comprehensive OpenAI provider logic for LLM features (located in src/utils/providers/model.ts), but the text-to-speech pipeline deliberately avoids OpenAI's TTS API. According to the source implementation, this architectural decision stems from several factors:
- Cost Efficiency – Edge TTS provides high-quality neural voices without requiring API key management or per-character billing associated with OpenAI's service.
- Voice Variety – Microsoft Edge TTS offers extensive language coverage with region-specific neural voices, enabling the per-language voice mapping that Read Frog implements.
- Offline Capability – While not fully offline, Edge TTS operates through the browser's edge services without external API authentication, simplifying the extension's permission model.
Consequently, users looking for OpenAI TTS integration will not find it in Read Frog; instead, they gain access to a flexible, configuration-driven TTS system powered by Edge TTS.
Summary
- Read Frog implements Text-to-Speech through Microsoft Edge TTS, not OpenAI's TTS API, utilizing the
useTextToSpeechhook insrc/hooks/use-text-to-speech.tsx. - Voice selection relies on automatic language detection mapping ISO-639-3 codes to specific voices via
ttsConfig.languageVoices, with fallback tottsConfig.defaultVoice. - Speech speed control (rate), pitch, and volume are managed through numeric ranges (–100 to +100) converted to signed strings by
toSignedValueand passed to the Edge TTS synthesis request. - The architecture uses off-screen documents in
src/entrypoints/background/tts-playback.tsto bypass CSP restrictions and enable audio playback in browser extensions. - Audio chunks are cached using React Query to prevent redundant synthesis of identical text segments.
Frequently Asked Questions
Does Read Frog use OpenAI Text-to-Speech?
No, Read Frog does not integrate OpenAI's Text-to-Speech API. According to the source code in mengxi-ream/read-frog, the extension exclusively uses Microsoft's Edge TTS service to generate speech audio. While the codebase includes OpenAI provider logic for LLM features, the TTS pipeline is deliberately built on Edge TTS to avoid additional API costs and authentication complexity.
How does Read Frog select voices for different languages?
Read Frog implements automatic voice selection through the resolveVoiceForText function in src/hooks/use-text-to-speech.tsx. The system first detects the text's language using the detectLanguage utility, then queries the ttsConfig.languageVoices mapping (defined in src/types/config/tts.ts) for an ISO-639-3 code match. If no specific voice is configured for the detected language, the system falls back to ttsConfig.defaultVoice.
Can I adjust the speech speed in Read Frog?
Yes, Read Frog provides granular control over speech speed through the rate parameter in TTSConfig. Users can set values ranging from –100 to +100, where negative values slow down speech and positive values accelerate it. The toSignedValue function converts these integers to percentage strings (e.g., "+20%" or "-10%") before sending them to the Edge TTS synthesis request in src/hooks/use-text-to-speech.tsx.
What is the maximum text length Read Frog can synthesize at once?
Read Frog automatically handles long text by chunking it into segments of ≤ 2KB UTF-8 bytes using the splitTextByUtf8Bytes utility in src/utils/server/edge-tts/chunk.ts. This ensures compliance with Edge TTS API limitations while allowing users to synthesize lengthy articles seamlessly. The UI displays progress through currentChunk and totalChunk counters provided by the useTextToSpeech hook.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →