# How to Configure ICE Servers for WebRTC Deployment Behind NAT

> Configure ICE servers for WebRTC behind NAT. Learn how to set the SPEECH_TO_SPEECH_ICE_SERVERS environment variable for seamless STUN/TURN server integration in your deployment.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-31

---

**To enable WebRTC connectivity through NAT and firewalls in the Hugging Face Speech-to-Speech repository, set the `SPEECH_TO_SPEECH_ICE_SERVERS` environment variable with a JSON array of STUN and TURN server configurations on the backend, and optionally configure `RTC_ICE_SERVERS` for the client side.**

The Speech-to-Speech repository provides a full-stack realtime engine that exposes the OpenAI Realtime API over **WebRTC**. When deploying this system behind NAT (Network Address Translation) or corporate firewalls, you must configure ICE (Interactive Connectivity Establishment) servers to facilitate peer-to-peer connections. This guide covers the specific environment variables, source code locations, and deployment steps required to configure ICE servers for WebRTC deployment behind NAT.

## Why ICE Servers Are Required for WebRTC Behind NAT

WebRTC establishes direct peer-to-peer connections between browsers and servers, but NAT devices and firewalls block direct incoming connections. **ICE servers**—specifically **STUN** (Session Traversal Utilities for NAT) and **TURN** (Traversal Using Relays around NAT) servers—solve this by discovering public IP addresses and relaying media traffic when direct paths fail.

Without proper ICE configuration, the `RTCPeerConnection` cannot establish connectivity in symmetric NAT environments, Docker containers with isolated UDP ports, or networks with strict firewall rules. The Speech-to-Speech repository handles this through environment-driven configuration parsed at runtime.

## Backend Configuration: SPEECH_TO_SPEECH_ICE_SERVERS

The server-side WebRTC session relies on the `SPEECH_TO_SPEECH_ICE_SERVERS` environment variable to construct the `RTCConfiguration` object passed to the peer connection.

### Environment Variable Format

Set `SPEECH_TO_SPEECH_ICE_SERVERS` to a JSON list of `RTCIceServer` dictionaries:

```json
[
  {"urls": "stun:stun.l.google.com:19302"},
  {"urls": "turn:turn.mycompany.com:3478", "username": "myuser", "credential": "mysecret"}
]

```

If this variable is missing or malformed, the system falls back to **aiortc’s defaults**, which include host candidates and Google’s public STUN server. For production deployments behind restrictive NAT, always provide a dedicated TURN server in this list.

### Source Code Implementation

In [[`src/speech_to_speech/api/openai_realtime/webrtc_session.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/webrtc_session.py)](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/webrtc_session.py#L50-L66), the `rtc_configuration_from_env()` function parses this variable:

```python

# From webrtc_session.py

def rtc_configuration_from_env():
    ice_servers = os.environ.get("SPEECH_TO_SPEECH_ICE_SERVERS")
    if ice_servers:
        config = json.loads(ice_servers)
        return RTCConfiguration(iceServers=config)
    return None  # Uses aiortc defaults

```

The `WebRTCSession` class inherits from `SessionTransport` and passes this configuration to the `RTCPeerConnection` during initialization, ensuring the SDP handshake includes the relay candidates necessary for NAT traversal.

## Client-Side Configuration: RTC_ICE_SERVERS

For the browser client (used in the demo application), the server parses the `RTC_ICE_SERVERS` environment variable and injects the configuration into the HTML/JS response.

In [[`demo/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/demo/server.py)](https://github.com/huggingface/speech-to-speech/blob/main/demo/server.py#L84-L101), the `_parse_ice_servers()` method handles this:

```python

# From demo/server.py

def _parse_ice_servers(self):
    rtc_ice_servers = os.environ.get("RTC_ICE_SERVERS", "[]")
    try:
        return json.loads(rtc_ice_servers)
    except json.JSONDecodeError:
        # Fallback: parse comma-separated URLs

        return [{"urls": url.strip()} for url in rtc_ice_servers.split(",")]

```

The client-side JavaScript in [[`demo/rtc/s2s-rtc-client.js`](https://github.com/huggingface/speech-to-speech/blob/main/demo/rtc/s2s-rtc-client.js)](https://github.com/huggingface/speech-to-speech/blob/main/demo/rtc/s2s-rtc-client.js) receives this configuration when creating the `RTCPeerConnection`, ensuring both sides of the WebRTC handshake use compatible ICE candidates.

## Step-by-Step Deployment Guide

Follow these steps to configure ICE servers for WebRTC deployment behind NAT:

1. **Provision a TURN server**  
   Deploy **coturn**, or use commercial services like Twilio or Xirsys. Ensure UDP traffic (and optionally TCP) is allowed on port 3478. Create credentials (username/password) or time-limited tokens depending on your provider.

2. **Set the backend ICE configuration**  
   Export the `SPEECH_TO_SPEECH_ICE_SERVERS` variable before starting the service:
   ```bash
   export SPEECH_TO_SPEECH_ICE_SERVERS='[
     {"urls":"stun:stun.l.google.com:19302"},
     {"urls":"turn:turn.mycompany.com:3478","username":"myuser","credential":"mysecret"}
   ]'
   ```

3. **Expose UDP ports**  
   If running in Docker, map the TURN port: `-p 3478:3478/udp`. For Kubernetes, expose UDP via a `NodePort` or `LoadBalancer` service.

4. **Configure the client environment** (optional)  
   If the browser needs specific ICE servers (e.g., when the demo server runs behind a different NAT):
   ```bash
   export RTC_ICE_SERVERS='[
     {"urls":"turn:turn.mycompany.com:3478","username":"myuser","credential":"mysecret"}
   ]'
   ```

5. **Start the Speech-to-Speech service**  
   Launch with realtime mode enabled:
   ```bash
   uv run speech-to-speech --mode realtime ...
   ```

   The `WebRTCSession` will instantiate `RTCPeerConnection` with your ICE configuration during the SDP exchange.

6. **Verify connectivity**  
   Open the browser developer console and inspect `icecandidate` events. Look for `candidateType: "relay"` in the generated candidates. If only `host` candidates appear, the client cannot reach the TURN server—check firewall rules and port accessibility.

## How It Works Under the Hood

The `WebRTCSession` class manages the WebRTC transport layer. When a client connects, the server:

- Calls `rtc_configuration_from_env()` to build the `RTCConfiguration`
- Passes this to `RTCPeerConnection` during the SDP offer/answer exchange
- Uses `PipelineAudioTrack` to stream RTP audio frames through the established ICE connection
- Maintains the `oai-events` data channel for JSON protocol messages over the same ICE route

This architecture ensures that once the ICE candidates are exchanged and the relay connection is established, the speech-to-speech pipeline operates transparently regardless of the underlying NAT topology.

## Summary

- **Backend configuration** relies on the `SPEECH_TO_SPEECH_ICE_SERVERS` environment variable, parsed by `rtc_configuration_from_env()` in [`webrtc_session.py`](https://github.com/huggingface/speech-to-speech/blob/main/webrtc_session.py)
- **Client configuration** uses `RTC_ICE_SERVERS`, parsed by `_parse_ice_servers()` in [`demo/server.py`](https://github.com/huggingface/speech-to-speech/blob/main/demo/server.py)
- **JSON format** supports multiple STUN/TURN entries with authentication credentials
- **TURN servers are mandatory** for symmetric NAT or Docker deployments where direct UDP connectivity fails
- **Verification** requires checking browser console for "relay" type ICE candidates

## Frequently Asked Questions

### What is the difference between STUN and TURN servers?

**STUN servers** reveal your public IP address to help establish direct peer-to-peer connections through NAT. **TURN servers** act as media relays when direct connections fail (such as with symmetric NAT or strict firewalls). For production Speech-to-Speech deployments, always include a TURN server in your `SPEECH_TO_SPEECH_ICE_SERVERS` configuration to ensure reliable connectivity.

### Why does the Speech-to-Speech repository use environment variables for ICE configuration?

Environment variables allow operators to configure ICE servers without modifying source code or rebuilding containers. This separation of concerns enables the same Docker image to work across different network environments—development, staging, and production—by simply changing the `SPEECH_TO_SPEECH_ICE_SERVERS` value at runtime.

### How do I verify that my TURN server is being used?

Open the browser's WebRTC internals (chrome://webrtc-internals/ in Chrome or the Firefox Web Console) and examine the ICE candidate statistics. If you see `candidateType: "relay"` with the IP address of your TURN server, the traffic is being relayed successfully. If you only see `host` or `srflx` (server reflexive) candidates, the TURN server is not being utilized.

### Can I use a commercial TURN service like Twilio or Xirsys?

Yes. Any STUN/TURN service providing RFC 5766-compliant servers works with the Speech-to-Speech repository. Simply format the service credentials as JSON objects in the `SPEECH_TO_SPEECH_ICE_SERVERS` array, including the `urls`, `username`, and `credential` fields provided by your commercial vendor.