How the Music Assistant Cache Controller Optimizes Metadata Caching

TLDR: The Cache Controller in the music-assistant/server repository optimizes metadata caching through SQLite-backed atomic storage, SHA-256 checksum validation, TTL-based expiration timestamps, and automatic LRU eviction based on database size limits, ensuring fast lookups without unbounded growth on embedded hardware.

The music-assistant/server repository implements a sophisticated metadata persistence layer to minimize redundant API calls and accelerate media lookups. At the heart of this system lies the Cache Controller (music_assistant/controllers/cache/controller.py), which employs multiple optimization strategies to balance performance, data integrity, and resource constraints on devices like the Raspberry Pi.

Atomic Storage with SQLite

Database-Backed Persistence

Unlike file-based caching systems that scatter data across directories, the controller utilizes a single SQLite database (cache.db) for all entries. As initialized in the CacheController constructor, this approach provides atomic write operations, ACID compliance, and indexed lookups without the overhead of a full database server. The compact storage format proves especially critical on resource-constrained hardware where memory and disk space are limited.

Data Integrity Through Checksums and Expiration

SHA-256 Validation

The _store_entry method implements corruption detection by calculating an SHA-256 checksum of the JSON-serialized value. This hash is stored alongside the data payload in the database, enabling the controller to validate entry integrity on every read. If a checksum mismatch occurs—indicating file corruption or disk errors—the controller automatically invalidates the entry and triggers a fresh fetch from the provider.

TTL-Based Expiration Tracking

Rather than performing expensive full-table scans to locate stale data, the controller stores an explicit expire_ts (Unix timestamp in seconds) with each entry. This timestamp allows the get method to reject expired data immediately during lookup, while the background cleanup process targets only obsolete records for efficient deletion.

Intelligent Cache Organization

Provider-Scoped Namespacing

To prevent key collisions between different music sources, the _make_key method constructs composite identifiers using the format <provider>|<category>|<key>. This architecture ensures that metadata from Spotify, Yandex, or local files remains isolated, enabling per-provider cache invalidation without affecting other data sources. The category parameter further segregates data types such as artists, albums, and playlists.

Persistent vs. Ephemeral Entries

The set method accepts a persistent boolean parameter that determines eviction priority. When persistent=True, entries are exempt from the size-based LRU eviction logic, ensuring that user-defined playlists, manual edits, and critical configuration survive automatic cleanup cycles while transient metadata (album art, temporary API responses) remains eligible for deletion.

Automated Resource Management

Size-Based LRU Eviction

The _auto_cleanup method monitors the total database size against the configurable MAX_CACHE_DB_SIZE_MB threshold. When the cache exceeds this limit, the controller executes a least-recently-used (LRU) deletion strategy, removing the oldest accessed rows until the storage footprint returns below the constraint. This guarantees predictable storage usage on embedded hardware.

Background Cleanup Daemon

A scheduled task (_auto_cleanup_task) executes every five minutes to remove expired entries and enforce size limits. This daemon operates independently of application logic, ensuring the database remains optimized without requiring explicit maintenance calls from other controllers or provider modules.

Flexible Access Patterns

For scenarios requiring real-time data—such as OAuth token refreshes or immediate playlist synchronization—the get method supports a bypass parameter. When bypass=True, the controller skips the cache lookup entirely, fetching fresh data directly from the provider while preserving the caching layer for standard operations.

Implementation Examples

The following examples demonstrate how to interact with the cache controller for typical metadata operations:


# Storing album metadata with 24-hour expiration

await mass.cache.set(
    key="album:42",
    value={"title": "The Dark Side of the Moon", "artist": "Pink Floyd"},
    category="album",
    persistent=False,
    ttl=86400,  # 24 hours in seconds

)

# Retrieving data with optional bypass for fresh fetch

album = await mass.cache.get(
    key="album:42",
    category="album",
    bypass=False,  # Set to True to skip cache

    default={},
)

For comprehensive API documentation and edge-case handling, refer to the test suite in tests/core/test_cache_controller.py.

Summary

  • SQLite Backend: Provides atomic writes and compact storage through cache.db without external database dependencies, as implemented in the CacheController constructor.
  • Checksum Validation: SHA-256 hashes in _store_entry detect corruption automatically, ensuring data integrity through stored checksums.
  • Expiration Tracking: Unix timestamps (expire_ts) enable instant stale data detection without full table scans.
  • Namespaced Isolation: The _make_key method prevents provider collisions using <provider>|<category>|<key> formatting.
  • Tiered Persistence: The persistent flag protects critical user data from LRU eviction while allowing temporary data cleanup.
  • Automated Maintenance: The _auto_cleanup_task daemon runs every five minutes to enforce MAX_CACHE_DB_SIZE_MB limits and remove expired entries via _auto_cleanup.
  • Selective Bypass: The bypass parameter in get enables real-time data fetching when cached values are unacceptable.

Frequently Asked Questions

How does the Cache Controller detect corrupted cache entries?

The controller calculates an SHA-256 checksum of the JSON-serialized value in _store_entry and stores it alongside the data. On retrieval, it validates the stored checksum against the current data; mismatches trigger automatic invalidation and fresh fetches from the provider.

What happens when the cache database grows too large?

When the database size exceeds MAX_CACHE_DB_SIZE_MB, the _auto_cleanup method executes an LRU eviction strategy, deleting the least-recently-used rows until the storage footprint returns below the threshold. This process runs automatically every five minutes via the _auto_cleanup_task background daemon.

Can specific metadata entries persist through automatic cleanups?

Yes. By setting persistent=True when calling set(), you mark entries as exempt from size-based LRU eviction. This is ideal for user-defined playlists and manual edits that must survive cache purges while temporary data gets removed.

How do I fetch fresh data without disabling the cache entirely?

Use the bypass=True parameter in the get method. This skips the cache lookup for that specific call, fetching directly from the provider, while leaving the caching mechanism active for all other operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →