# How Telegram-iOS Processes Voice Messages and Generates Audio Waveforms: The AudioWaveform Module Explained

> Discover how Telegram-iOS processes voice messages, generating real-time audio waveforms via the AudioWaveform module. Learn about PCM capture, Opus compression, and efficient bitstream encoding.

- Repository: [TelegramMessenger/Telegram-iOS](https://github.com/TelegramMessenger/Telegram-iOS)
- Tags: internals
- Published: 2026-04-07

---

**Telegram-iOS captures raw PCM audio via a custom Audio Unit recorder, compresses it to Opus, and generates a real-time 100-sample waveform preview that is encoded into a compact 5-bit bitstream for efficient transport and rendering.**

The `TelegramMessenger/Telegram-iOS` repository implements a sophisticated audio pipeline for voice messages that balances real-time performance with bandwidth efficiency. At the heart of this system lies the **AudioWaveform** module, which transforms raw audio samples into compact visual representations suitable for messaging. This article examines the complete data flow from low-level PCM capture to UI rendering.

## Low-Level Audio Capture with ManagedAudioRecorder

### RemoteIO Audio Unit and PCM Sampling

In [`submodules/TelegramUI/Sources/ManagedAudioRecorder.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/TelegramUI/Sources/ManagedAudioRecorder.swift), the `ManagedAudioRecorder` class initializes a low-level **RemoteIO Audio Unit** to capture uncompressed PCM samples as `Int16` arrays. This approach provides direct access to audio buffers without the latency of high-level frameworks like `AVAudioRecorder`.

The recorder implements `ManagedAudioRecorderContext.processWaveformPreview`, which processes incoming sample buffers in real time:

```swift
let recorder = ManagedAudioRecorderImpl(...)
recorder.start()

// Inside ManagedAudioRecorderContext.processWaveformPreview(...)
for i in 0..<count {
    var sample = samples.advanced(by: i).pointee
    sample = sample < 0 ? (sample == Int16.min ? Int16.max : -sample) : sample
    // Track peak...
    // Every N samples store the peak in compressedWaveformSamples
}

```

### OGG Opus Compression

Parallel to buffering, the recorder streams samples to `TGOggOpusWriter` for real-time compression. The Opus encoder runs continuously, ensuring that when recording stops, the compressed audio file is already written to disk while the waveform preview remains available in memory.

## Real-Time Waveform Generation

### Peak Detection and Compression Strategy

During capture, the recorder monitors incoming PCM samples to build a **preview waveform**. Each `Int16` value is normalized to its absolute magnitude, tracking the current peak across a configurable `peakCompressionFactor` sample window. When the window closes, the peak value is appended to `compressedWaveformSamples`.

This peak-detection approach discards temporal resolution in favor of amplitude accuracy, ensuring the visual waveform represents the audio's dynamic range even during extended recordings.

### Dynamic Scaling Algorithm

To maintain a fixed preview size regardless of recording duration, Telegram-iOS implements dynamic downscaling. Once `compressedWaveformSamples` reaches **200 entries**, the buffer is halved to **100 peaks**, and the compression factor doubles. This recursive decimation preserves the waveform's overall shape while bounding memory usage, preventing unbounded growth during long voice messages.

## The AudioWaveform Data Structure

### Encoding to 5-Bit Bitstreams

When recording completes, the `takeData()` method finalizes the waveform. The 100-sample preview is normalized to a peak value of **31** (hard-coded to match the 5-bit display requirement) and clamped to the range 0–31. The `AudioWaveform` struct in [`submodules/AudioWaveform/Sources/AudioWaveform.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/AudioWaveform/Sources/AudioWaveform.swift) encapsulates this data:

```swift
let scaledSamplesMemory = malloc(100 * 2)!
let scaledSamples = scaledSamplesMemory.assumingMemoryBound(to: Int16.self)
defer { free(scaledSamplesMemory) }

memset(scaledSamples, 0, 100 * 2)
// …fill `scaledSamples` with the 100‑sample preview…

let waveform = AudioWaveform(samples: Data(bytes: scaledSamplesMemory,
                                            count: 100 * 2), peak: 31)
let bitstream = waveform.makeBitstream()           // 5‑bit per sample

```

The `makeBitstream()` method encodes the array into a compact binary format: `(samples/2 * 5)/8` bytes plus a 4-byte header. This **5-bit per sample** encoding achieves a 62.5% size reduction compared to raw 16-bit integers.

### Decoding and Sub-Waveform Extraction

On the receiving side, `init(bitstream:bitsPerSample:)` reconstructs the 100-sample array from the bitstream:

```swift
let receivedBitstream: Data = …               // from server / local storage
let waveform = AudioWaveform(bitstream: receivedBitstream,
                             bitsPerSample: 5)

```

The `subwaveform(from:to:)` method enables partial rendering for playback progress indicators, returning a sliced `AudioWaveform` without decompressing the entire dataset:

```swift
let sub = waveform.subwaveform(from: 0.0, to: playbackProgress)
let scrubber = AudioWaveformComponent(samples: sub.samples,
                                      peak: sub.peak,
                                      color: .gray,
                                      progress: playbackProgress,
                                      trimRange: nil)

```

## UI Rendering Components

### AudioWaveformComponent for ComponentKit

The `AudioWaveformComponent` in [`submodules/TelegramUI/Components/AudioWaveformComponent/Sources/AudioWaveformComponent.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/TelegramUI/Components/AudioWaveformComponent/Sources/AudioWaveformComponent.swift) renders waveforms as vertical bar charts. It accepts `samples` and `peak` parameters, calculating bar heights proportional to the 5-bit amplitude values:

```swift
let component = AudioWaveformComponent(
    samples: waveform.samples,
    peak: waveform.peak,
    color: .white,
    progress: 0.0,
    trimRange: nil)
let view = component.makeView()

```

### AudioWaveformNode for AsyncDisplayKit

For message bubbles and preview panels, `AudioWaveformNode` (in [`submodules/TelegramUI/Components/AudioWaveformNode/Sources/AudioWaveformNode.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/TelegramUI/Components/AudioWaveformNode/Sources/AudioWaveformNode.swift)) provides an `ASDisplayNode` implementation that draws waveforms directly using Core Animation. This node is instantiated in [`submodules/TelegramUI/Sources/ChatControllerMediaRecording.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/TelegramUI/Sources/ChatControllerMediaRecording.swift) to display the recording preview in chat interfaces.

The `RecordedAudioData` struct produced by `ManagedAudioRecorder` carries both the compressed Opus file and the waveform bitstream, ensuring the visual representation travels with the audio data through the Telegram network.

## Summary

- **PCM Capture**: `ManagedAudioRecorder` uses a RemoteIO Audio Unit to capture raw `Int16` samples with minimal latency.
- **Opus Compression**: Audio streams to `TGOggOpusWriter` simultaneously with waveform generation, producing platform-standard voice message files.
- **Peak Tracking**: Real-time waveform generation uses absolute-value peak detection and dynamic compression (200-to-100 sample reduction) to maintain fixed memory footprints.
- **Compact Encoding**: The `AudioWaveform` type encodes 100 samples into a 5-bit bitstream, achieving significant bandwidth savings over uncompressed audio data.
- **Dual Rendering**: The system provides both `AudioWaveformComponent` (ComponentKit) and `AudioWaveformNode` (AsyncDisplayKit) for flexible UI integration across chat bubbles, preview panels, and scrubbing interfaces.

## Frequently Asked Questions

### How does Telegram-iOS compress audio waveforms for transmission?

Telegram-iOS compresses waveforms using a **5-bit per sample** bitstream encoding implemented in `AudioWaveform.makeBitstream()`. The 100-sample waveform is normalized to a peak value of 31 (fitting within 5 bits), then packed into `(samples/2 * 5)/8` bytes plus a 4-byte header. This reduces the waveform payload from 200 bytes (raw Int16) to approximately 67 bytes, making it efficient for real-time messaging.

### What bit depth does Telegram use for voice message waveforms?

The waveform visualization uses **5-bit resolution** per sample (values 0–31), hard-coded in the `AudioWaveform` peak property. This limitation aligns with the UI design, which renders 5-bit amplitude bars, while the underlying audio recording maintains full 16-bit PCM fidelity for the actual Opus-encoded voice file.

### How does the waveform preview update during recording?

During recording, `ManagedAudioRecorderContext` tracks the absolute peak of incoming PCM samples over configurable intervals. After every 200 compressed peaks, the buffer scales down to 100 peaks by averaging pairs, doubling the compression factor. This dynamic scaling ensures the preview size remains constant regardless of recording duration, updating the visual representation in real time without reallocating buffers.

### Where is the AudioWaveform generated in the source code?

The `AudioWaveform` object is generated in [`submodules/TelegramUI/Sources/ManagedAudioRecorder.swift`](https://github.com/TelegramMessenger/Telegram-iOS/blob/main/submodules/TelegramUI/Sources/ManagedAudioRecorder.swift) within the `takeData()` method of `ManagedAudioRecorderContext`. This method finalizes the 100-sample preview, normalizes it to peak 31, and constructs the `AudioWaveform` instance that gets embedded into `RecordedAudioData` for transmission alongside the Opus-compressed audio file.