How WebRTC Is Integrated for Voice Calls in Telegram-iOS
Telegram-iOS integrates WebRTC through a custom Objective-C wrapper named TgVoipWebrtc that bridges iOS AVAudioSession with the WebRTC C++ library, implementing a bespoke audio device module to handle real-time voice packet exchange.
The Telegram-iOS client utilizes the TgVoipWebrtc submodule to enable peer-to-peer voice communication. This architecture wraps the official WebRTC C++ library in Objective-C glue code, creating a custom audio transport that mediates between iOS native audio sessions and WebRTC's media processing pipeline. The integration handles everything from audio session configuration to DTLS-SRTP encryption negotiation while exposing a clean Swift API for the application layer.
Step-by-Step Integration Flow
The WebRTC integration follows a six-stage pipeline that moves from audio hardware initialization to encrypted media exchange.
Creating the WebRTC Audio Device
In SharedCallAudioDevice.mm, the system constructs an AudioDeviceModuleIOS (webrtc::tgcalls_ios_adm::AudioDeviceModuleIOS) and encapsulates it within a WrappedAudioDeviceModuleIOS. This custom module implements the webrtc::AudioTransport interface and serves as the bridge between iOS RTCAudioSession and the WebRTC native layer. The file submodules/TgVoipWebrtc/Sources/SharedCallAudioDevice.mm contains the factory logic that instantiates this device with specific recording and system mute capabilities.
Configuring the iOS Audio Session
Before initializing WebRTC, SharedCallAudioDevice.setupAudioSession() establishes the iOS audio environment. This method creates an RTCAudioSessionConfiguration object configured with AVAudioSessionModeVoiceChat, enabling Bluetooth A2DP routes and optimizing buffer sizes for real-time communication. The configuration activates the RTCAudioSession singleton, ensuring the app captures microphone input and routes output to the appropriate speaker or headset.
Initializing the WebRTC Instance
The OngoingCallThreadLocalContextWebrtc class in OngoingCallThreadLocalContext.mm registers concrete tgcalls::Instance implementations (including InstanceImpl and InstanceV2Impl) and constructs a tgcalls::Instance. This initialization accepts the previously built audio module, optional video capturers, and signaling callback closures. At line 1520 of OngoingCallThreadLocalContext.mm, the context creates the native instance that will manage the DTLS handshake and SRTP key exchange.
Registering Audio Callbacks
The WrappedAudioDeviceModuleIOS maintains a registry of webrtc::AudioTransport callbacks. When WebRTC requires PCM data, it invokes RecordedDataIsAvailable for microphone input or NeedMorePlayData for speaker output. These methods forward audio buffers between the WebRTC media engine and the iOS RTCAudioSession audio queue, handling the actual voice packet capture and playback. This bidirectional flow operates on line 138 of OngoingCallThreadLocalContext.mm within the WrappedAudioDeviceModuleIOS implementation.
Starting the Call
From the Swift layer, OngoingCallContext.swift instantiates OngoingCallThreadLocalContextWebrtc and supplies a signaling closure (sendSignalingData) for exchanging ICE candidates and session descriptions with Telegram's servers. Calling start() on the context initiates the WebRTC negotiation sequence, establishes the DTLS/SRTP media session, and begins streaming audio frames through the custom device module configured in the previous steps.
Tear-Down and Resource Cleanup
When a call terminates, SharedCallAudioDevice invokes ActualStop() on the WrappedAudioDeviceModuleIOS instance. This stops the playout and recording threads, releases the RTCAudioSession, and deallocates the audio buffers. The Swift layer simultaneously calls context.stop() to signal the media engine and disable the manual audio session activation.
Code Implementation Examples
The following snippets demonstrate the concrete implementation patterns used throughout the codebase.
Initializing the VoIP Engine in Swift
Before starting a call, the application creates the audio device and thread-local context:
let audioDevice = SharedCallAudioDevice(disableRecording: false,
enableSystemMute: true)
SharedCallAudioDevice.setupAudioSession()
let context = OngoingCallThreadLocalContextWebrtc(
queue: OngoingCallThreadLocalContextQueueImpl(queue: DispatchQueue.main),
sendSignalingData: { data in
// forward `data` to Telegram signalling server
},
videoCapturer: nil) // voice‑only call
Starting a Voice Call
Once initialized, the context begins the media session with connection parameters and encryption keys:
context.start(
connection: OngoingCallConnectionDescription(
connectionId: 12345,
ip: "185.76.10.2",
ipv6: "",
port: 443,
peerTag: Data()),
encryptionKey: Data(), // DTLS‑SRTP key
isOutgoing: true,
isVideoEnabled: false,
videoCapture: nil,
audioDevice: audioDevice)
Audio Callback Flow in C++
The following simplified C++ illustrates how the wrapper forwards audio data between WebRTC and iOS audio queues:
// WebRTC asks for recorded mic data
int32_t WrappedAudioDeviceModuleIOS::RecordedDataIsAvailable(
const void* audioSamples, size_t nSamples, …) {
_mutex.Lock();
for (auto &pair : _audioTransports) {
if (pair.first) {
pair.first->RecordedDataIsAvailable(
audioSamples, nSamples, …);
}
}
_mutex.Unlock();
return 0;
}
Here, pair.first represents the RTCAudioSession transport that ultimately writes PCM data into the system audio queue.
Stopping the Call
Call termination requires coordinated cleanup between Swift and the underlying modules:
context.stop { reason, callId, ... in
// clean‑up UI
}
audioDevice.setManualAudioSessionIsActive(false) // disables RTCAudioSession
Key Source Files and Components
The WebRTC integration spans several files across the TgVoipWebrtc and TelegramVoip submodules:
submodules/TgVoipWebrtc/Sources/OngoingCallThreadLocalContext.mm— Objective-C wrapper around WebRTCInstanceand audio modules that exposes the core call logic to Swift.submodules/TgVoipWebrtc/PublicHeaders/TgVoipWebrtc/OngoingCallThreadLocalContext.h— Public header defining the Swift-compatible API for call contexts.submodules/TgVoipWebrtc/Sources/SharedCallAudioDevice.mm— Implementation of the custom iOS audio device that conforms towebrtc::AudioTransport.submodules/TelegramVoip/Sources/OngoingCallContext.swift— Swift façade that orchestrates call creation, signaling, and UI integration.submodules/TgVoipWebrtc/Sources/OngoingCallThreadLocalContextVideoCapturer.mm— Video capture bridge for video calls (unused in voice-only scenarios).
Summary
- Telegram-iOS embeds WebRTC via the
TgVoipWebrtcsubmodule, which wraps the C++ WebRTC library in Objective-C. - A custom
AudioDeviceModuleIOSbridges iOSAVAudioSessionwith WebRTC's media engine through theWrappedAudioDeviceModuleIOSclass. - Call initialization involves configuring the audio session, instantiating
tgcalls::Instance, and registering audio transport callbacks inSharedCallAudioDevice.mm. - The Swift layer in
OngoingCallContext.swiftdrives the call lifecycle, supplying signaling data and managing DTLS-SRTP encryption parameters. - Audio data flows bidirectionally between the iOS system and WebRTC via
RecordedDataIsAvailableandNeedMorePlayDatacallbacks.
Frequently Asked Questions
How does Telegram-iOS handle audio routing between WebRTC and iOS system audio?
Telegram-iOS implements a custom WrappedAudioDeviceModuleIOS that implements the webrtc::AudioTransport interface. This module registers itself with the WebRTC Instance and forwards audio buffers to the iOS RTCAudioSession. When WebRTC requests microphone data via RecordedDataIsAvailable, the wrapper pulls PCM samples from the iOS audio queue; conversely, NeedMorePlayData pushes decoded audio to the speaker. This architecture allows WebRTC to remain platform-agnostic while the wrapper handles iOS-specific audio session management.
What is the role of the TgVoipWebrtc submodule in voice calls?
The TgVoipWebrtc submodule serves as the integration layer between Telegram's UI code and the WebRTC native library. It contains Objective-C++ glue code that constructs the AudioDeviceModuleIOS, manages RTCAudioSession configuration with voice-chat optimizations, and exposes safe Swift bindings through OngoingCallThreadLocalContext. This submodule encapsulates all platform-specific audio and video capture logic required by WebRTC.
Where does the signaling data flow in the WebRTC integration?
Signaling data flows through the sendSignalingData closure supplied during OngoingCallThreadLocalContextWebrtc initialization in OngoingCallContext.swift. This Swift closure receives ICE candidates and session descriptions from the WebRTC native layer and forwards them to Telegram's servers. The return path delivers remote signaling data back to the WebRTC instance via the context's processing methods, completing the handshake required for DTLS-SRTP encryption negotiation.
How is the iOS audio session configured for voice calls?
The SharedCallAudioDevice.setupAudioSession() method configures the RTCAudioSession with AVAudioSessionModeVoiceChat, enabling Bluetooth A2DP routing and optimizing buffer sizes for low-latency communication. This configuration activates the audio session before the WebRTC instance starts and deactivates it via setManualAudioSessionIsActive(false) after the call ends, ensuring proper audio resource management and preventing conflicts with other iOS audio apps.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →