# How the Muse Audio Export Pipeline Converts PCM to WAV and MP3

> Learn how the Muse audio export pipeline converts PCM to WAV and MP3. Discover the use of LAME encoder and silence padding for efficient audio processing.

- Repository: [Ko Shin/muse](https://github.com/kkoshin/muse)
- Tags: internals
- Published: 2026-03-05

---

**The Muse audio export pipeline merges raw PCM audio segments into a temporary WAV file with configurable silence padding, then transcodes it to MP3 using a platform-specific LAME encoder.**

The **audio export pipeline** in [kkoshin/muse](https://github.com/kkoshin/muse) handles the complete journey from raw audio fragments to distributable MP3 files. According to the source code, the pipeline operates in two distinct phases: first constructing a valid RIFF/WAV container from multiple PCM sources, then re-encoding that WAV stream into a compressed MP3 with embedded ID3 metadata.

## Pipeline Initialization and Configuration

The entry point for audio export is the `rememberAudioExportPipeline` composable function, which constructs an `AudioExportPipeline` instance in [`AudioExportPipeline.kt`](https://github.com/kkoshin/muse/blob/main/AudioExportPipeline.kt). This class stores the input PCM file paths, an optional list of textual phrases (used to calculate dynamic silence durations), a `paddingSilence` configuration, and default mono `AudioSampleMetadata`.

When the pipeline starts via the `start` method, it accepts a `BufferedSink` (typically writing to a file in the app's cache directory) where the final MP3 data will be streamed. The pipeline manages its own temporary file lifecycle, creating a transient WAV file during processing that is deleted automatically after successful MP3 conversion.

## Merging PCM Segments into a WAV Container

Inside `generateMp3Data`, the pipeline writes each input PCM file sequentially to a temporary sink. After every audio segment, the function calls `getSilence` to append padding based on the `SilenceDuration` configuration:

- **Fixed duration**: A constant silence length (e.g., 2 seconds) between all clips
- **Dynamic duration**: Calculated from the length of the corresponding phrase text, allowing variable pauses based on content context

Once all PCM data and silence padding are written to the temporary file, `WaveHeaderWriter` generates a proper RIFF/WAV header by updating the file with correct byte counts and format specifications. This produces a standards-compliant WAV file that intermediate parsers can read.

## Transcoding WAV to MP3 with ID3 Metadata

After the WAV header is finalized, the pipeline invokes `encodeWavAsMp3` to begin the second phase. The `WavParser` class reads PCM samples from the temporary WAV file, streaming them into a platform-specific `Mp3Encoder` implementation.

The encoder writes MP3 frames directly to the output sink supplied by the caller. During this process, it injects an `Mp3Metadata` object containing ID3 tags—specifically setting the artist field to "μ's" and the year to the current calendar year—before the first audio frame is encoded.

## Platform-Specific LAME Implementations

The MP3 encoding logic uses Kotlin Multiplatform's `expect`/`actual` mechanism to wrap native LAME libraries on each platform.

### iOS Implementation

In [`Mp3Encoder.ios.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.ios.kt), the encoder interfaces with the native *lame* library via Kotlin/Native C interop. The implementation configures the sample rate and channel count, then repeatedly calls `lame_encode_buffer` (or `lame_encode_buffer_interleaved` for stereo) to convert PCM buffers into MP3 frames. Finally, `lame_encode_flush` writes any remaining internal encoder data to the sink.

### Android Implementation

In [`Mp3Encoder.android.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.android.kt), the pipeline wraps the Java *android-lame* library. It instantiates an `AndroidLame` object with the appropriate bitrate and channel mode, then feeds mono or stereo PCM buffers to `AndroidLame.encode`. The `AndroidLame.flush` method ensures all delayed MP3 data is written to complete the file.

## Usage Examples

### Exporting with Jetpack Compose

```kotlin
val pcmFiles = listOf(
    Path("audio1.pcm"),
    Path("audio2.pcm")
)

val exportPipeline = rememberAudioExportPipeline(
    input = pcmFiles,
    paddingSilence = 2.seconds   // add 2 s of silence between clips
)

LaunchedEffect(Unit) {
    val exportFile = SystemFileSystem.createCacheFile("final.mp3")
    SystemFileSystem.sink(exportFile).buffer().use { sink ->
        exportPipeline.start(sink)
    }
}

```

### Manual Pipeline Invocation

```kotlin
suspend fun exportAudio(pcmPaths: List<Path>, outMp3: Path) {
    val pipeline = AudioExportPipeline(
        pcmInputs = pcmPaths,
        phrases = emptyList(),
        paddingSilence = SilenceDuration.Fixed(1.seconds)
    )
    SystemFileSystem.sink(outMp3).buffer().use { sink ->
        pipeline.start(sink)
    }
}

```

### Inspecting Intermediate WAV Data

```kotlin
val wavTmp = createCacheFile("temp.wav", false)
SystemFileSystem.sink(wavTmp).buffer().use { sink ->
    AudioExportPipeline(pcmInputs, emptyList(), SilenceDuration.Fixed(0.seconds))
        .generateMp3Data(sink)
}
WaveHeaderWriter(wavTmp, MonoAudioSampleMetadata()).writeHeader()
println("WAV length ≈ ${WavParser(SystemFileSystem.source(wavTmp).buffer()).getLengthSeconds()} sec")

```

## Summary

- The **audio export pipeline** in [`AudioExportPipeline.kt`](https://github.com/kkoshin/muse/blob/main/AudioExportPipeline.kt) orchestrates a two-stage conversion: PCM merging followed by MP3 transcoding.
- Raw PCM files are concatenated with configurable **silence padding** (fixed or phrase-length dynamic) into a temporary WAV file using `WaveHeaderWriter`.
- **Platform-specific LAME encoders** handle the WAV-to-MP3 conversion, with [`Mp3Encoder.ios.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.ios.kt) using native interop and [`Mp3Encoder.android.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.android.kt) using the `android-lame` Java library.
- The pipeline embeds ID3 metadata (artist "μ's" and current year) during the encoding phase and reports progress throughout both stages.
- Temporary WAV files are automatically cleaned up after successful MP3 generation, leaving only the final compressed output.

## Frequently Asked Questions

### How does the pipeline determine the duration of silence between audio segments?

The pipeline checks the `SilenceDuration` type passed to `AudioExportPipeline`. If `Fixed`, it uses a constant duration value. If `Dynamic`, it calculates silence length based on the character count or estimated reading time of the corresponding phrase text provided in the constructor, allowing pauses that match speech rhythm.

### Why does the pipeline create an intermediate WAV file instead of encoding MP3 directly from PCM?

The pipeline constructs a temporary WAV container to ensure **audio sample alignment** and **header validation** before MP3 encoding. The `WaveHeaderWriter` guarantees that concatenated PCM fragments form a valid RIFF structure that `WavParser` can reliably read, abstracting away raw byte stream complexities from the LAME encoder implementations.

### What ID3 metadata does the exported MP3 contain?

According to the `encodeWavAsMp3` implementation in [`AudioExportPipeline.kt`](https://github.com/kkoshin/muse/blob/main/AudioExportPipeline.kt), the pipeline injects an `Mp3Metadata` object setting the **artist tag to "μ's"** (the Greek letter mu with an apostrophe) and the **year tag to the current system year** at export time.

### Can the pipeline handle stereo audio or only mono?

While the default configuration uses mono `AudioSampleMetadata`, the platform-specific encoders in both [`Mp3Encoder.ios.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.ios.kt) and [`Mp3Encoder.android.kt`](https://github.com/kkoshin/muse/blob/main/Mp3Encoder.android.kt) support stereo output. The iOS implementation calls `lame_encode_buffer_interleaved` for stereo streams, while the Android implementation configures `AndroidLame` with the appropriate channel mode based on the input WAV format parsed by `WavParser`.