How the Muse Audio Export Pipeline Converts PCM to WAV and MP3
The Muse audio export pipeline merges raw PCM audio segments into a temporary WAV file with configurable silence padding, then transcodes it to MP3 using a platform-specific LAME encoder.
The audio export pipeline in kkoshin/muse handles the complete journey from raw audio fragments to distributable MP3 files. According to the source code, the pipeline operates in two distinct phases: first constructing a valid RIFF/WAV container from multiple PCM sources, then re-encoding that WAV stream into a compressed MP3 with embedded ID3 metadata.
Pipeline Initialization and Configuration
The entry point for audio export is the rememberAudioExportPipeline composable function, which constructs an AudioExportPipeline instance in AudioExportPipeline.kt. This class stores the input PCM file paths, an optional list of textual phrases (used to calculate dynamic silence durations), a paddingSilence configuration, and default mono AudioSampleMetadata.
When the pipeline starts via the start method, it accepts a BufferedSink (typically writing to a file in the app's cache directory) where the final MP3 data will be streamed. The pipeline manages its own temporary file lifecycle, creating a transient WAV file during processing that is deleted automatically after successful MP3 conversion.
Merging PCM Segments into a WAV Container
Inside generateMp3Data, the pipeline writes each input PCM file sequentially to a temporary sink. After every audio segment, the function calls getSilence to append padding based on the SilenceDuration configuration:
- Fixed duration: A constant silence length (e.g., 2 seconds) between all clips
- Dynamic duration: Calculated from the length of the corresponding phrase text, allowing variable pauses based on content context
Once all PCM data and silence padding are written to the temporary file, WaveHeaderWriter generates a proper RIFF/WAV header by updating the file with correct byte counts and format specifications. This produces a standards-compliant WAV file that intermediate parsers can read.
Transcoding WAV to MP3 with ID3 Metadata
After the WAV header is finalized, the pipeline invokes encodeWavAsMp3 to begin the second phase. The WavParser class reads PCM samples from the temporary WAV file, streaming them into a platform-specific Mp3Encoder implementation.
The encoder writes MP3 frames directly to the output sink supplied by the caller. During this process, it injects an Mp3Metadata object containing ID3 tags—specifically setting the artist field to "μ's" and the year to the current calendar year—before the first audio frame is encoded.
Platform-Specific LAME Implementations
The MP3 encoding logic uses Kotlin Multiplatform's expect/actual mechanism to wrap native LAME libraries on each platform.
iOS Implementation
In Mp3Encoder.ios.kt, the encoder interfaces with the native lame library via Kotlin/Native C interop. The implementation configures the sample rate and channel count, then repeatedly calls lame_encode_buffer (or lame_encode_buffer_interleaved for stereo) to convert PCM buffers into MP3 frames. Finally, lame_encode_flush writes any remaining internal encoder data to the sink.
Android Implementation
In Mp3Encoder.android.kt, the pipeline wraps the Java android-lame library. It instantiates an AndroidLame object with the appropriate bitrate and channel mode, then feeds mono or stereo PCM buffers to AndroidLame.encode. The AndroidLame.flush method ensures all delayed MP3 data is written to complete the file.
Usage Examples
Exporting with Jetpack Compose
val pcmFiles = listOf(
Path("audio1.pcm"),
Path("audio2.pcm")
)
val exportPipeline = rememberAudioExportPipeline(
input = pcmFiles,
paddingSilence = 2.seconds // add 2 s of silence between clips
)
LaunchedEffect(Unit) {
val exportFile = SystemFileSystem.createCacheFile("final.mp3")
SystemFileSystem.sink(exportFile).buffer().use { sink ->
exportPipeline.start(sink)
}
}
Manual Pipeline Invocation
suspend fun exportAudio(pcmPaths: List<Path>, outMp3: Path) {
val pipeline = AudioExportPipeline(
pcmInputs = pcmPaths,
phrases = emptyList(),
paddingSilence = SilenceDuration.Fixed(1.seconds)
)
SystemFileSystem.sink(outMp3).buffer().use { sink ->
pipeline.start(sink)
}
}
Inspecting Intermediate WAV Data
val wavTmp = createCacheFile("temp.wav", false)
SystemFileSystem.sink(wavTmp).buffer().use { sink ->
AudioExportPipeline(pcmInputs, emptyList(), SilenceDuration.Fixed(0.seconds))
.generateMp3Data(sink)
}
WaveHeaderWriter(wavTmp, MonoAudioSampleMetadata()).writeHeader()
println("WAV length ≈ ${WavParser(SystemFileSystem.source(wavTmp).buffer()).getLengthSeconds()} sec")
Summary
- The audio export pipeline in
AudioExportPipeline.ktorchestrates a two-stage conversion: PCM merging followed by MP3 transcoding. - Raw PCM files are concatenated with configurable silence padding (fixed or phrase-length dynamic) into a temporary WAV file using
WaveHeaderWriter. - Platform-specific LAME encoders handle the WAV-to-MP3 conversion, with
Mp3Encoder.ios.ktusing native interop andMp3Encoder.android.ktusing theandroid-lameJava library. - The pipeline embeds ID3 metadata (artist "μ's" and current year) during the encoding phase and reports progress throughout both stages.
- Temporary WAV files are automatically cleaned up after successful MP3 generation, leaving only the final compressed output.
Frequently Asked Questions
How does the pipeline determine the duration of silence between audio segments?
The pipeline checks the SilenceDuration type passed to AudioExportPipeline. If Fixed, it uses a constant duration value. If Dynamic, it calculates silence length based on the character count or estimated reading time of the corresponding phrase text provided in the constructor, allowing pauses that match speech rhythm.
Why does the pipeline create an intermediate WAV file instead of encoding MP3 directly from PCM?
The pipeline constructs a temporary WAV container to ensure audio sample alignment and header validation before MP3 encoding. The WaveHeaderWriter guarantees that concatenated PCM fragments form a valid RIFF structure that WavParser can reliably read, abstracting away raw byte stream complexities from the LAME encoder implementations.
What ID3 metadata does the exported MP3 contain?
According to the encodeWavAsMp3 implementation in AudioExportPipeline.kt, the pipeline injects an Mp3Metadata object setting the artist tag to "μ's" (the Greek letter mu with an apostrophe) and the year tag to the current system year at export time.
Can the pipeline handle stereo audio or only mono?
While the default configuration uses mono AudioSampleMetadata, the platform-specific encoders in both Mp3Encoder.ios.kt and Mp3Encoder.android.kt support stereo output. The iOS implementation calls lame_encode_buffer_interleaved for stereo streams, while the Android implementation configures AndroidLame with the appropriate channel mode based on the input WAV format parsed by WavParser.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →