Read Frog Text-to-Speech Implementation: Edge TTS Voice Selection and Speed Control

Read Frog does not integrate OpenAI's Text-to-Speech API; instead, it implements TTS through Microsoft Edge TTS with automatic language-based voice selection and granular control over speech rate, pitch, and volume.

While many modern reading applications integrate OpenAI's Text-to-Speech service for neural voice generation, the open-source Read Frog browser extension takes a different architectural approach. According to the source code in mengxi-ream/read-frog, the project deliberately avoids OpenAI TTS in favor of Microsoft's Edge TTS API, which provides robust multilingual support without requiring separate API subscriptions. This implementation offers sophisticated voice selection logic based on automatic language detection and fine-grained audio controls.

Architecture Overview

The TTS pipeline in Read Frog operates through a six-stage process orchestrated by the useTextToSpeech React hook in src/hooks/use-text-to-speech.tsx:

  1. Trigger – UI components invoke the play method with text content and TTSConfig.
  2. Voice Resolution – resolveVoiceForText detects the input language via detectLanguage and selects a voice from ttsConfig.languageVoices, falling back to ttsConfig.defaultVoice if no language-specific mapping exists.
  3. Chunking – splitTextByUtf8Bytes divides text into chunks ≤ 2KB UTF-8 bytes to respect API limits.
  4. Synthesis Request – Each chunk triggers sendMessage("edgeTtsSynthesize", …) with payload parameters including voice, rate, pitch, and volume.
  5. Audio Playback – The background script src/entrypoints/background/tts-playback.ts creates an off-screen document to bypass Content Security Policy restrictions and forwards base64 audio via ttsOffscreenPlay.
  6. Caching – React Query caches generated audio chunks under ["tts-audio", …] keys to prevent redundant synthesis.

Voice Selection Logic

Read Frog implements data-driven voice selection through the TTSConfig type defined in src/types/config/tts.ts. Rather than relying on OpenAI's limited voice roster, the system maps ISO-639-3 language codes to specific Edge TTS voices:

  • Language Detection – The detectLanguage utility analyzes input text to determine its language code.
  • Voice Mapping – The system queries ttsConfig.languageVoices using the detected code. If a mapping exists (e.g., "eng" → "en-US-AriaNeural"), that voice is selected.
  • Fallback Strategy – When no language-specific voice is configured or detection fails, the system defaults to ttsConfig.defaultVoice.

This approach allows users to configure distinct voices for different languages, enabling multilingual reading sessions with appropriate native pronunciation.

Speech Speed and Audio Controls

While OpenAI TTS offers limited speed control (typically 0.25x to 4.0x multipliers), Read Frog provides granular adjustment of three audio parameters through the TTSConfig interface:

  • Rate (speech speed): Range –100 to +100, converted to percentage strings (e.g., "+10%" or "-20%")
  • Pitch: Range –100 to +100, converted to Hertz strings (e.g., "+5Hz")
  • Volume: Range –100 to +100, converted to percentage strings

The toSignedValue utility function in src/hooks/use-text-to-speech.tsx handles the transformation:

function toSignedValue(value: number, unit: "%" | "Hz"): string {
  return `${value >= 0 ? "+" : ""}${value}${unit}`
}

These values are passed to the Edge TTS API through the sendMessage("edgeTtsSynthesize", …) payload, allowing real-time adjustment of speech characteristics without modifying the underlying audio files.

Code Implementation Examples

React Hook Usage

import { useTextToSpeech } from "@/hooks/use-text-to-speech"
import { useConfig } from "@/hooks/use-config"

function SpeakButton({ text }: { text: string }) {
  const { config } = useConfig()
  const { play, stop, isPlaying, currentChunk, totalChunks } = useTextToSpeech()

  const handleClick = async () => {
    await play(text, config.tts)   // config.tts implements TTSConfig
  }

  return (
    <button onClick={handleClick} disabled={isPlaying}>
      {isPlaying ? `Playing ${currentChunk}/${totalChunks}` : "Speak"}
    </button>
  )
}

Source: src/hooks/use-text-to-speech.tsx

Voice Resolution Implementation

async function resolveVoiceForText(text: string, ttsConfig: TTSConfig): Promise<string> {
  const detectedLanguage = await detectLanguage(text, {
    minLength: 0,
    enableLLM: ttsConfig.detectLanguageMode === "llm",
  })

  if (detectedLanguage && detectedLanguage in ttsConfig.languageVoices) {
    return ttsConfig.languageVoices[detectedLanguage as keyof typeof ttsConfig.languageVoices] ?? ttsConfig.defaultVoice
  }
  return ttsConfig.defaultVoice
}

Source: src/hooks/use-text-to-speech.tsx

Edge TTS Synthesis Request

const response = await sendMessage("edgeTtsSynthesize", {
  text: chunk,
  voice,
  rate: toSignedValue(ttsConfig.rate, "%"),
  pitch: toSignedValue(ttsConfig.pitch, "Hz"),
  volume: toSignedValue(ttsConfig.volume, "%"),
})

Source: src/hooks/use-text-to-speech.tsx

Why Read Frog Uses Edge TTS Instead of OpenAI

The Read Frog codebase contains comprehensive OpenAI provider logic for LLM features (located in src/utils/providers/model.ts), but the text-to-speech pipeline deliberately avoids OpenAI's TTS API. According to the source implementation, this architectural decision stems from several factors:

  • Cost Efficiency – Edge TTS provides high-quality neural voices without requiring API key management or per-character billing associated with OpenAI's service.
  • Voice Variety – Microsoft Edge TTS offers extensive language coverage with region-specific neural voices, enabling the per-language voice mapping that Read Frog implements.
  • Offline Capability – While not fully offline, Edge TTS operates through the browser's edge services without external API authentication, simplifying the extension's permission model.

Consequently, users looking for OpenAI TTS integration will not find it in Read Frog; instead, they gain access to a flexible, configuration-driven TTS system powered by Edge TTS.

Summary

  • Read Frog implements Text-to-Speech through Microsoft Edge TTS, not OpenAI's TTS API, utilizing the useTextToSpeech hook in src/hooks/use-text-to-speech.tsx.
  • Voice selection relies on automatic language detection mapping ISO-639-3 codes to specific voices via ttsConfig.languageVoices, with fallback to ttsConfig.defaultVoice.
  • Speech speed control (rate), pitch, and volume are managed through numeric ranges (–100 to +100) converted to signed strings by toSignedValue and passed to the Edge TTS synthesis request.
  • The architecture uses off-screen documents in src/entrypoints/background/tts-playback.ts to bypass CSP restrictions and enable audio playback in browser extensions.
  • Audio chunks are cached using React Query to prevent redundant synthesis of identical text segments.

Frequently Asked Questions

Does Read Frog use OpenAI Text-to-Speech?

No, Read Frog does not integrate OpenAI's Text-to-Speech API. According to the source code in mengxi-ream/read-frog, the extension exclusively uses Microsoft's Edge TTS service to generate speech audio. While the codebase includes OpenAI provider logic for LLM features, the TTS pipeline is deliberately built on Edge TTS to avoid additional API costs and authentication complexity.

How does Read Frog select voices for different languages?

Read Frog implements automatic voice selection through the resolveVoiceForText function in src/hooks/use-text-to-speech.tsx. The system first detects the text's language using the detectLanguage utility, then queries the ttsConfig.languageVoices mapping (defined in src/types/config/tts.ts) for an ISO-639-3 code match. If no specific voice is configured for the detected language, the system falls back to ttsConfig.defaultVoice.

Can I adjust the speech speed in Read Frog?

Yes, Read Frog provides granular control over speech speed through the rate parameter in TTSConfig. Users can set values ranging from –100 to +100, where negative values slow down speech and positive values accelerate it. The toSignedValue function converts these integers to percentage strings (e.g., "+20%" or "-10%") before sending them to the Edge TTS synthesis request in src/hooks/use-text-to-speech.tsx.

What is the maximum text length Read Frog can synthesize at once?

Read Frog automatically handles long text by chunking it into segments of ≤ 2KB UTF-8 bytes using the splitTextByUtf8Bytes utility in src/utils/server/edge-tts/chunk.ts. This ensures compliance with Edge TTS API limitations while allowing users to synthesize lengthy articles seamlessly. The UI displays progress through currentChunk and totalChunk counters provided by the useTextToSpeech hook.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →