Voice Data and Subscription Quota Caching Strategies in Muse
Muse implements a three-tier caching architecture for voice data—in-memory voice lists, on-disk generated audio via DataStore, and a dedicated iOS cache directory—while subscription quotas remain uncached and are fetched in real-time from the TTS provider.
The open-source Muse repository (kkoshin/muse) provides a Kotlin Multiplatform framework for text-to-speech processing that balances performance optimization with accurate resource tracking. Understanding how the library handles caching voice data and subscription quotas reveals a deliberate architectural choice: aggressive local caching for expensive voice generation operations alongside direct API queries for billing-sensitive quota information.
Three-Tier Voice Data Caching Architecture
Muse employs a multi-layered caching strategy to minimize redundant network requests and audio regeneration. The system operates across memory, persistent preferences, and the file system to maximize reuse of voice assets.
In-Memory Voice List Caching
The SpeechProcessorManager maintains an in-memory cache to avoid repeated network calls when fetching available voices. According to the implementation in muse/src/iosMain/kotlin/io/github/kkoshin/muse/core/manager/SpeechProcessorManager.ios.kt (lines 48‑56), the class stores a voicesCache property holding the List<Voice> returned by TTSProvider.queryVoices().
When calling queryVoiceList(skipCache: Boolean), setting skipCache = true forces a fresh network request, while the default false returns the cached list if available. This pattern ensures rapid UI responsiveness during voice selection screens while allowing manual refresh when needed.
On-Disk Generated Audio Caching
For generated audio files, Muse uses a DataStore<Preferences> named voicePreference to map text-voice combinations to their corresponding file system paths. As implemented in SpeechProcessorManager.ios.kt (lines 73‑89 and 99‑108), the cache key follows the format "{voiceId}_{text}" for short text, or a hash digest for long text inputs.
The getOrGenerate() method first queries this DataStore, then verifies the file exists at the returned URI before deciding whether to generate new audio. This prevents redundant text-to-speech processing for identical inputs, significantly reducing API costs and latency for repetitive content.
File System Cache Directory
The library uses MusePathManager to resolve a dedicated voices subdirectory within the iOS Caches directory (NSCachesDirectory). As defined in muse/src/iosMain/kotlin/io/github/kkoshin/muse/repo/MusePathManager.ios.kt (lines 20‑30), this folder is created on demand and serves as the permanent storage location for downloaded voice assets and generated audio files, ensuring data persists across app launches while remaining eligible for system cleanup if storage runs low.
Real-Time Subscription Quota Management
Unlike voice data, subscription quotas are intentionally not cached. The system queries quota information directly from the TTS provider whenever the UI requires it, such as in the Settings screen.
Direct Quota Queries
The SpeechProcessorManager.queryQuota() method (line 46 in SpeechProcessorManager.ios.kt) forwards requests directly to TTSProvider.queryQuota() without any local caching layer. The returned CharacterQuota data class—defined in muse/src/commonMain/kotlin/io/github/kkoshin/muse/core/provider/TTSProvider.kt (lines 35‑42)—contains consumed, total, and status fields, with a computed remaining property derived from these values.
This real-time approach ensures users always see current usage limits and prevents the app from consuming quota based on stale data, which is critical for billing-sensitive operations.
Practical Implementation Examples
Retrieving a Cached Voice List
// Force refresh by setting skipCache = true
val result = speechProcessorManager.queryVoiceList(skipCache = false)
result.onSuccess { voiceList ->
// voiceList comes from voicesCache if previously loaded
}
Generating or Reusing Audio for Short Text
val pathResult = speechProcessorManager.getOrGenerate(
voiceId = "en_us_001",
text = "Hello world"
)
pathResult.onSuccess { audioPath ->
// audioPath points to locally stored file, either cached or freshly generated
}
Handling Long Text with Hash-Based Keys
val longResult = speechProcessorManager.getOrGenerateForLongText(
voiceId = "en_us_001",
longText = veryLongString
)
longResult.onSuccess { audioPath ->
// Reuses cached MP3 if hash exists, otherwise generates new file
}
Querying Current Subscription Usage
val quotaResult = speechProcessorManager.queryQuota()
quotaResult.onSuccess { quota ->
println("Remaining characters: ${quota.remaining}")
// Always reflects real-time provider data
}
Summary
- Triple-layer caching: Muse combines in-memory voice lists (
voicesCache), DataStore-backed audio file mappings (voicePreference), and iOS Cache directory storage for comprehensive voice data reuse. - Cache bypass capability: The
skipCacheparameter inqueryVoiceList()allows forced network refresh when fresh data is required. - Intelligent key generation: Short text uses
"{voiceId}_{text}"keys while long text utilizes hash-based identifiers to handle arbitrary-length content efficiently. - No quota caching:
CharacterQuotadata is always fetched live fromTTSProviderto ensure billing accuracy and prevent over-limit usage. - File verification: The
getOrGenerate()method validates file existence on disk before returning cached paths, preventing errors from stale DataStore entries.
Frequently Asked Questions
How does Muse prevent redundant audio generation for identical text inputs?
Muse prevents redundant generation by using a DataStore<Preferences> named voicePreference that maps unique voice-text combinations to file system paths. When getOrGenerate() is called, it constructs a cache key (either "{voiceId}_{text}" for short strings or a hash for long text), checks the DataStore for an existing path, and verifies the file exists before returning the cached audio or generating new content.
Why doesn't Muse cache subscription quota data locally?
Subscription quotas are never cached because accurate real-time usage tracking is critical for billing and rate limiting. The SpeechProcessorManager forwards queryQuota() calls directly to the TTSProvider without intermediate storage, ensuring the CharacterQuota object reflects actual current limits and consumption as maintained by the backend service.
What happens if a cached audio file is deleted outside the app?
The getOrGenerate() method in SpeechProcessorManager includes a file existence check after retrieving the path from voicePreference. If the file no longer exists on disk (whether deleted by the system during low storage cleanup or manual removal), the method proceeds to generate new audio and updates the DataStore with the new path, maintaining cache integrity automatically.
Can the voice list cache be cleared programmatically?
Yes, the voice list cache can be bypassed by passing skipCache = true to queryVoiceList(), which forces a fresh network request to TTSProvider.queryVoices(). While there is no explicit "clear cache" method shown in the source, setting skipCache ensures fresh data retrieval, and the voicesCache property would be overwritten with the new results on successful network responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →