How Telegram-iOS Processes Voice Messages and Generates Audio Waveforms: The AudioWaveform Module Explained
Telegram-iOS captures raw PCM audio via a custom Audio Unit recorder, compresses it to Opus, and generates a real-time 100-sample waveform preview that is encoded into a compact 5-bit bitstream for efficient transport and rendering.
The TelegramMessenger/Telegram-iOS repository implements a sophisticated audio pipeline for voice messages that balances real-time performance with bandwidth efficiency. At the heart of this system lies the AudioWaveform module, which transforms raw audio samples into compact visual representations suitable for messaging. This article examines the complete data flow from low-level PCM capture to UI rendering.
Low-Level Audio Capture with ManagedAudioRecorder
RemoteIO Audio Unit and PCM Sampling
In submodules/TelegramUI/Sources/ManagedAudioRecorder.swift, the ManagedAudioRecorder class initializes a low-level RemoteIO Audio Unit to capture uncompressed PCM samples as Int16 arrays. This approach provides direct access to audio buffers without the latency of high-level frameworks like AVAudioRecorder.
The recorder implements ManagedAudioRecorderContext.processWaveformPreview, which processes incoming sample buffers in real time:
let recorder = ManagedAudioRecorderImpl(...)
recorder.start()
// Inside ManagedAudioRecorderContext.processWaveformPreview(...)
for i in 0..<count {
var sample = samples.advanced(by: i).pointee
sample = sample < 0 ? (sample == Int16.min ? Int16.max : -sample) : sample
// Track peak...
// Every N samples store the peak in compressedWaveformSamples
}
OGG Opus Compression
Parallel to buffering, the recorder streams samples to TGOggOpusWriter for real-time compression. The Opus encoder runs continuously, ensuring that when recording stops, the compressed audio file is already written to disk while the waveform preview remains available in memory.
Real-Time Waveform Generation
Peak Detection and Compression Strategy
During capture, the recorder monitors incoming PCM samples to build a preview waveform. Each Int16 value is normalized to its absolute magnitude, tracking the current peak across a configurable peakCompressionFactor sample window. When the window closes, the peak value is appended to compressedWaveformSamples.
This peak-detection approach discards temporal resolution in favor of amplitude accuracy, ensuring the visual waveform represents the audio's dynamic range even during extended recordings.
Dynamic Scaling Algorithm
To maintain a fixed preview size regardless of recording duration, Telegram-iOS implements dynamic downscaling. Once compressedWaveformSamples reaches 200 entries, the buffer is halved to 100 peaks, and the compression factor doubles. This recursive decimation preserves the waveform's overall shape while bounding memory usage, preventing unbounded growth during long voice messages.
The AudioWaveform Data Structure
Encoding to 5-Bit Bitstreams
When recording completes, the takeData() method finalizes the waveform. The 100-sample preview is normalized to a peak value of 31 (hard-coded to match the 5-bit display requirement) and clamped to the range 0–31. The AudioWaveform struct in submodules/AudioWaveform/Sources/AudioWaveform.swift encapsulates this data:
let scaledSamplesMemory = malloc(100 * 2)!
let scaledSamples = scaledSamplesMemory.assumingMemoryBound(to: Int16.self)
defer { free(scaledSamplesMemory) }
memset(scaledSamples, 0, 100 * 2)
// …fill `scaledSamples` with the 100‑sample preview…
let waveform = AudioWaveform(samples: Data(bytes: scaledSamplesMemory,
count: 100 * 2), peak: 31)
let bitstream = waveform.makeBitstream() // 5‑bit per sample
The makeBitstream() method encodes the array into a compact binary format: (samples/2 * 5)/8 bytes plus a 4-byte header. This 5-bit per sample encoding achieves a 62.5% size reduction compared to raw 16-bit integers.
Decoding and Sub-Waveform Extraction
On the receiving side, init(bitstream:bitsPerSample:) reconstructs the 100-sample array from the bitstream:
let receivedBitstream: Data = … // from server / local storage
let waveform = AudioWaveform(bitstream: receivedBitstream,
bitsPerSample: 5)
The subwaveform(from:to:) method enables partial rendering for playback progress indicators, returning a sliced AudioWaveform without decompressing the entire dataset:
let sub = waveform.subwaveform(from: 0.0, to: playbackProgress)
let scrubber = AudioWaveformComponent(samples: sub.samples,
peak: sub.peak,
color: .gray,
progress: playbackProgress,
trimRange: nil)
UI Rendering Components
AudioWaveformComponent for ComponentKit
The AudioWaveformComponent in submodules/TelegramUI/Components/AudioWaveformComponent/Sources/AudioWaveformComponent.swift renders waveforms as vertical bar charts. It accepts samples and peak parameters, calculating bar heights proportional to the 5-bit amplitude values:
let component = AudioWaveformComponent(
samples: waveform.samples,
peak: waveform.peak,
color: .white,
progress: 0.0,
trimRange: nil)
let view = component.makeView()
AudioWaveformNode for AsyncDisplayKit
For message bubbles and preview panels, AudioWaveformNode (in submodules/TelegramUI/Components/AudioWaveformNode/Sources/AudioWaveformNode.swift) provides an ASDisplayNode implementation that draws waveforms directly using Core Animation. This node is instantiated in submodules/TelegramUI/Sources/ChatControllerMediaRecording.swift to display the recording preview in chat interfaces.
The RecordedAudioData struct produced by ManagedAudioRecorder carries both the compressed Opus file and the waveform bitstream, ensuring the visual representation travels with the audio data through the Telegram network.
Summary
- PCM Capture:
ManagedAudioRecorderuses a RemoteIO Audio Unit to capture rawInt16samples with minimal latency. - Opus Compression: Audio streams to
TGOggOpusWritersimultaneously with waveform generation, producing platform-standard voice message files. - Peak Tracking: Real-time waveform generation uses absolute-value peak detection and dynamic compression (200-to-100 sample reduction) to maintain fixed memory footprints.
- Compact Encoding: The
AudioWaveformtype encodes 100 samples into a 5-bit bitstream, achieving significant bandwidth savings over uncompressed audio data. - Dual Rendering: The system provides both
AudioWaveformComponent(ComponentKit) andAudioWaveformNode(AsyncDisplayKit) for flexible UI integration across chat bubbles, preview panels, and scrubbing interfaces.
Frequently Asked Questions
How does Telegram-iOS compress audio waveforms for transmission?
Telegram-iOS compresses waveforms using a 5-bit per sample bitstream encoding implemented in AudioWaveform.makeBitstream(). The 100-sample waveform is normalized to a peak value of 31 (fitting within 5 bits), then packed into (samples/2 * 5)/8 bytes plus a 4-byte header. This reduces the waveform payload from 200 bytes (raw Int16) to approximately 67 bytes, making it efficient for real-time messaging.
What bit depth does Telegram use for voice message waveforms?
The waveform visualization uses 5-bit resolution per sample (values 0–31), hard-coded in the AudioWaveform peak property. This limitation aligns with the UI design, which renders 5-bit amplitude bars, while the underlying audio recording maintains full 16-bit PCM fidelity for the actual Opus-encoded voice file.
How does the waveform preview update during recording?
During recording, ManagedAudioRecorderContext tracks the absolute peak of incoming PCM samples over configurable intervals. After every 200 compressed peaks, the buffer scales down to 100 peaks by averaging pairs, doubling the compression factor. This dynamic scaling ensures the preview size remains constant regardless of recording duration, updating the visual representation in real time without reallocating buffers.
Where is the AudioWaveform generated in the source code?
The AudioWaveform object is generated in submodules/TelegramUI/Sources/ManagedAudioRecorder.swift within the takeData() method of ManagedAudioRecorderContext. This method finalizes the 100-sample preview, normalizes it to peak 31, and constructs the AudioWaveform instance that gets embedded into RecordedAudioData for transmission alongside the Opus-compressed audio file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →