Music Assistant HTTP Streaming Pipeline: Codecs, Sample Rates, and Transcoding Explained
Music Assistant normalizes every HTTP audio stream to PCM format, automatically detecting actual sample rates from the source and resampling to match player-supported rates, ensuring seamless playback across diverse codecs and hardware capabilities.
The music-assistant/server repository implements a sophisticated audio pipeline that handles HTTP streaming from providers like Spotify, Tidal, and internet radio. This architecture abstracts codec complexity by converting all incoming streams—including FLAC, AAC, MP3, and OGG—to a uniform raw PCM format while respecting player-specific sample rate limitations.
How Providers Supply Source Codec Information
Every music provider implements a get_stream_details() method that returns a StreamDetails object containing the source audio_format. This data structure, defined in music_assistant_models.media_items.AudioFormat, captures the codec (content_type), sample rate, bit depth, and channel count of the original stream.
In music_assistant/providers/yandex_music/streaming.py and similar provider implementations, the returned object looks like this:
sd = StreamDetails(
path="https://example.com/track.flac",
stream_type=StreamType.HTTP,
audio_format=AudioFormat(
content_type=ContentType.FLAC,
sample_rate=44100,
bit_depth=16,
channels=2
),
)
The StreamsAudio controller in music_assistant/controllers/streams/audio.py consumes these details to determine whether the stream can pass through directly or requires transcoding.
HTTP Stream Pipeline Architecture
Provider to StreamDetails Resolution
When a playback request arrives, StreamsAudio.get_stream_details() extracts the URL and validates the audio format. For internet radio streams (ICY, HLS, or plain HTTP), the system calls resolve_radio_stream() to handle redirects, parse HLS manifests, and convert ICY streams into direct URLs before proceeding.
Raw vs. Transcode Decision Logic
The pipeline applies a strict rule for codec handling:
- Direct streaming: Only if the source is already raw PCM (
ContentType.PCM) and the player supports the exact sample rate does the framework skip FFmpeg and stream bytes directly. - Transcoding: For all other codecs (MP3, AAC, FLAC, OGG), the pipeline always routes through FFmpeg to produce a raw PCM stream that the internal mixer can process.
This guarantees that the SmartFadesMixer receives consistent audio data regardless of source format variability.
FFmpeg Integration and Sample Rate Detection
Building the Conversion Arguments
The FFMpeg class in music_assistant/helpers/ffmpeg.py constructs command-line arguments using get_ffmpeg_args(), mapping the provider's input_format to the player's required output_format:
ffmpeg_args = get_ffmpeg_args(
input_format=input_format,
output_format=output_format,
filter_params=filter_params or [],
extra_args=extra_args or [],
input_path=audio_input if isinstance(audio_input, str) else "-",
output_path=audio_output if isinstance(audio_output, str) else "-",
loglevel=loglevel,
)
By default, FFmpeg outputs PCM s16le (signed 16-bit little-endian), though the system can request alternative bit depths based on player capabilities.
Runtime Sample Rate Detection
While FFmpeg processes the stream, the _log_reader_task parses stderr output using regex patterns (_FFMPEG_SAMPLE_RATE_RE, _FFMPEG_BIT_RATE_RE) defined at the top of the file. The detected actual sample rate—which may differ from the provider's initial hint—is stored in self.input_stream_info and mirrored back to the shared AudioFormat instance.
This ensures downstream components work with the true sample rate rather than estimated metadata:
ffmpeg = FFMpeg(
audio_input=details.path,
input_format=details.audio_format,
output_format=AudioFormat(
content_type=ContentType.PCM,
sample_rate=target_rate,
bit_depth=16,
channels=2
),
)
await ffmpeg.start()
# Contains the real detected sample rate from the source
print(ffmpeg.input_stream_info.sample_rate)
Sample Rate Normalization for Player Compatibility
Supported Rate Snapping
Players expose their hardware limitations through a list of supported sample rates (e.g., [44100, 48000, 96000]). The audio.py controller provides two helper functions to map desired rates to viable hardware options:
_snap_supported_rate_up: Selects the smallest supported rate greater than or equal to the target, defaulting to the maximum available._snap_supported_rate_down: Selects the largest supported rate less than or equal to the target, defaulting to the minimum available.
def _snap_supported_rate_up(target: int, supported_sample_rates: list[int]) -> int:
"""Pick the smallest supported rate ≥ target, otherwise return max."""
valid_rates = [r for r in supported_sample_rates if r >= target]
return min(valid_rates) if valid_rates else max(supported_sample_rates)
def _snap_supported_rate_down(target: int, supported_sample_rates: list[int]) -> int:
"""Pick the largest supported rate ≤ target, otherwise return min."""
valid_rates = [r for r in supported_sample_rates if r <= target]
return max(valid_rates) if valid_rates else min(supported_sample_rates)
Resampling Strategy
If the player requests a specific output sample rate (or OUTPUT_AUTO selects the highest supported rate) that differs from the FFmpeg-detected input rate, the pipeline instructs FFmpeg to resample using asetrate or aresample filters. This ensures the final PCM stream matches exactly what the hardware expects, preventing playback errors or audio artifacts.
Complete Pipeline Implementation Example
The following pattern demonstrates how the controllers orchestrate the full flow from provider to player:
# 1. Retrieve source details from provider
details = await mass.streams.get_stream_details(queue_item, seek_position=0)
# 2. Check actual detected sample rate after FFmpeg parsing
actual_rate = details.audio_format.sample_rate # e.g., 44100
# 3. Determine target rate based on player capabilities
player_cfg = {
"output_sample_rate": 48000, # or "auto"
"supported_sample_rates": [44100, 48000, 96000]
}
target_rate = player_cfg["output_sample_rate"]
if target_rate == "auto":
target_rate = max(player_cfg["supported_sample_rates"])
# 4. FFmpeg handles transcoding and resampling in one step
# The output PCM stream is now normalized for the specific player
Summary
- Source abstraction: Providers declare codec and sample rate via
StreamDetails.audio_format, but the actual values are validated by FFmpeg at runtime. - Universal PCM output: All non-PCM codecs (MP3, AAC, FLAC, OGG) transit through FFmpeg to produce raw PCM streams compatible with the internal mixer.
- Dynamic rate detection: FFmpeg stderr parsing updates
input_stream_infowith the true sample rate, correcting mismatches between metadata and reality. - Hardware adaptation: The
_snap_supported_rate_*helpers ensure output rates always match player-supported values, with automatic resampling when necessary. - Architecture path: Provider →
StreamsAudio→FFMpegwrapper →SmartFadesMixer→ Player.
Frequently Asked Questions
Does Music Assistant always transcode audio files?
No. The system only transcodes when necessary. If the source is already raw PCM (ContentType.PCM) and the sample rate matches the player's capabilities exactly, Music Assistant streams the bytes directly without invoking FFmpeg. For all compressed formats (MP3, AAC, FLAC, OGG), transcoding to PCM is mandatory to ensure the internal mixer can apply crossfades and volume normalization.
How does Music Assistant handle internet radio streams with variable bit rates?
For radio streams (ICY, HLS, or HTTP), StreamsAudio in music_assistant/controllers/streams/audio.py calls resolve_radio_stream() to resolve redirects and normalize the URL. Once the stream begins, FFmpeg parses the actual audio parameters from the stream headers in real-time. The detected sample rate and codec from input_stream_info override any initial guesses, ensuring the pipeline adapts to changes in stream quality or format.
What happens if my audio player only supports 48 kHz but the source is 44.1 kHz?
The pipeline automatically resamples. Using the _snap_supported_rate_up and _snap_supported_rate_down helpers in audio.py, the system selects the closest supported rate (in this case, 48000 Hz if supported, or 44100 Hz if not). FFmpeg then applies resampling filters (aresample or asetrate) to convert the PCM stream to the target rate before sending it to the player.
Can I force Music Assistant to output a specific sample rate?
Yes. Players can configure output_sample_rate to a specific integer (e.g., 48000) or set it to "auto" to use the highest supported rate. The StreamsAudio controller reads this configuration and passes the target rate to the FFmpeg wrapper as the output_format.sample_rate parameter, overriding the automatic detection logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →