# Voice Burst Packet Protocol in BitChat: Implementation and Processing Flow

> Explore BitChat's Voice Burst Packet Protocol for live push-to-talk voice. Learn about its implementation, packet types, and how ChatLiveVoiceCoordinator assembles playable audio.

- Repository: [permissionlesstech/bitchat](https://github.com/permissionlesstech/bitchat)
- Tags: internals
- Published: 2026-08-08

---

**BitChat implements a lightweight custom protocol for transmitting live push-to-talk voice bursts over its mesh network, using `VoiceBurstPacket` objects with an 8-byte burstID header, sequence numbers, and flag-based packet types (START, DATA, END, CANCELED) that are assembled by `ChatLiveVoiceCoordinator` into playable audio files.**

BitChat's mesh networking stack supports real-time voice communication through a specialized voice burst packet protocol. This system fragments continuous audio into discrete `VoiceBurstPacket` instances that traverse the network inside Noise session payloads (`NoisePayloadType.voiceFrame`) or public mesh broadcasts (`MessageType.voiceFrame`), enabling low-latency push-to-talk functionality across BLE and mesh transports.

## Voice Burst Packet Structure and Wire Format

### Header and Payload Layout

Each voice burst packet follows a strict binary wire format defined in [[`bitchat/Protocols/VoiceBurstPacket.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Protocols/VoiceBurstPacket.swift)](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Protocols/VoiceBurstPacket.swift):

```

[burstID: 8 bytes][seq: UInt16 BE][flags: UInt8][payload…]

```

The **burstID** is an 8-byte cryptographically secure random identifier generated via `VoiceBurstPacket.makeBurstID`. The **seq** field uses big-endian UInt16 encoding, starting at 0 for the START packet and incrementing for each subsequent DATA packet. The END packet receives its own sequence number in the series.

### Packet Types and Flag Semantics

The protocol defines four distinct packet types controlled by the flags byte:

- **`0x01` (START)** – Initiates a new burst. Payload contains `codec: UInt8` identifying the audio codec (e.g., AAC LC 16kHz mono).
- **`0x00` (DATA)** – Carries one or more encoded audio frames. Payload format repeats `[length: UInt16 BE][frame data]` for each frame.
- **`0x02` (END)** – Terminates the burst. Payload contains `totalDataPackets: UInt16 BE` followed by `durationMs: UInt32 BE`.
- **`0x04` (CANCELED)** – Aborts an in-progress burst. Contains no payload; signals the receiver to discard the assembly.

## Encoding and Decoding Voice Burst Packets

### Creating Voice Burst Packets

Outgoing voice bursts require constructing a sequence of packets with incrementing sequence numbers. The burst begins with a START packet containing codec information, followed by DATA packets carrying AAC frames, and concludes with an END packet containing metadata.

```swift
import Bitchat

let burstID = VoiceBurstPacket.makeBurstID()

// Create START packet (seq 0)
let start = VoiceBurstPacket(burstID: burstID,
                             seq: 0,
                             kind: .start(codec: .aacLC16kMono))!

// Create DATA packet (seq 1) with encoded frames
let data = VoiceBurstPacket(burstID: burstID,
                            seq: 1,
                            kind: .frames([frameData]))!

// Create END packet (seq 2) with summary statistics
let end = VoiceBurstPacket(burstID: burstID,
                           seq: 2,
                           kind: .end(totalDataPackets: 1,
                                      durationMs: 128))!

let encodedStart = start.encode()
let encodedData  = data.encode()
let encodedEnd   = end.encode()

```

### Decoding Inbound Payloads

Received payloads are deserialized using the static `decode` method. Failed decodings result in packet drops with logged warnings.

```swift
if let packet = VoiceBurstPacket.decode(incomingPayload) {
    // packet.kind indicates .start, .frames, .end, or .canceled
    switch packet.kind {
    case .start(let codec):
        // Initialize assembly with codec info
    case .frames(let frames):
        // Write frames to temporary buffer
    default:
        break
    }
}

```

### Packetizing Outgoing Audio Streams

For real-time capture, `VoiceBurstPacketizer` manages the fragmentation of continuous audio into maximum-size payloads respecting `TransportConfig.pttMaxBurstContentBytes`.

```swift
var packetizer = VoiceBurstPacketizer(burstID: burstID,
                                      budget: TransportConfig.pttMaxBurstContentBytes)

// Add frames as they become available
let flushed = packetizer.add(encodedAACFrame)
send(flushed)  // Immediately transmit any full packets

// Flush remaining data when burst completes
let final = packetizer.flush()
send(final)

// Transmit END packet with accurate statistics
let endPacket = VoiceBurstPacket(burstID: burstID,
                                 seq: packetizer.nextSeq,
                                 kind: .end(totalDataPackets: packetizer.dataPacketCount,
                                            durationMs: totalDurationMs))!
send([endPacket.encode()])

```

## Processing Pipeline in ChatLiveVoiceCoordinator

### Initial Packet Decoding

Inbound voice frames enter the system through [[`bitchat/ViewModels/ChatLiveVoiceCoordinator.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/ViewModels/ChatLiveVoiceCoordinator.swift)](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/ViewModels/ChatLiveVoiceCoordinator.swift). The `handle(_:from:scope:nickname:timestamp:)` method receives raw payloads and delegates deserialization to `VoiceBurstPacket.decode(_:)`.

### Assembly Management and State Tracking

The coordinator maintains active assemblies in a dictionary keyed by `AssemblyKey`, which uniquely identifies a burst using the composite of **peerID**, **scope**, and **burstID**. This prevents collision between simultaneous bursts from different peers or different conversation contexts.

When a START or stray DATA packet arrives without an existing assembly, `makeAssembly` creates a new entry in the coordinator's `assemblies` dictionary. Subsequent packets are routed via `apply(packet, to:)` to the correct assembly instance.

### Handling START and DATA Packets

START packets initialize the assembly with codec configuration and prepare a temporary "live-capture" file on disk. DATA packets extract individual AAC frames—up to `VoiceBurstPacket.maxFramesPerPacket` (8) per packet—and append them to this temporary file. A `PTTBurstPlayer` instance streams decoded frames to the UI in real-time during reception.

### Finalizing and Canceling Bursts

The END packet triggers assembly finalization: the temporary live file (e.g., `voice_live_… .m4a`) is closed, renamed to a permanent voice-note filename (`voice_… .m4a`), and the conversation UI replaces the temporary bubble with the finalized message containing the complete audio file.

CANCELED packets initiate cleanup: the coordinator discards the assembly record and deletes any partially written temporary files, preventing orphaned data from incomplete transmissions.

## Security and Resource Constraints

### Cryptographic Identifiers

The 8-byte **burstID** is generated using cryptographically secure random number generation via `VoiceBurstPacket.makeBurstID`, ensuring global uniqueness across the mesh network without coordination.

### Transport Budget and Concurrency Limits

The protocol enforces strict resource boundaries defined in [[`bitchat/Services/MeshTransportCapabilities.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Services/MeshTransportCapabilities.swift)](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Services/MeshTransportCapabilities.swift):

- **BLE Encryption Budget**: Each packet's payload respects `TransportConfig.pttMaxBurstContentBytes` to fit within single BLE frame constraints.
- **Concurrent Assemblies**: The system caps simultaneous burst processing via `TransportConfig.pttMaxConcurrentAssemblies` to prevent memory exhaustion.
- **Frame Packing**: Maximum 8 frames per DATA packet (`VoiceBurstPacket.maxFramesPerPacket`) optimizes the balance between latency and packet overhead.

### Sender Authentication

Assembly keys incorporate the authenticated **peerID** bound to the Noise session or cryptographically signed message. This prevents hijacking—a START packet from a different peer cannot assume an existing assembly because the peerID component of the key would mismatch. The scope differentiation ensures private direct messages and public mesh broadcasts maintain separate assembly namespaces.

## Summary

- **Wire Format**: Voice burst packets use an 8-byte burstID, UInt16 big-endian sequence numbers, and UInt8 flags to define START, DATA, END, and CANCELED packet types.
- **Assembly Management**: `ChatLiveVoiceCoordinator` tracks partial bursts using composite keys of peerID, scope, and burstID, preventing collisions and ensuring orderly reconstruction.
- **Real-Time Streaming**: Incoming DATA packets write AAC frames to temporary files that stream to the UI via `PTTBurstPlayer`, with END packets triggering finalization to permanent storage.
- **Security**: Burst IDs are cryptographically random and bound to authenticated sender identities; assemblies cannot be hijacked across peer boundaries.
- **Resource Constraints**: BLE frame budgets limit packet sizes, while `maxFramesPerPacket` (8) and `pttMaxConcurrentAssemblies` prevent memory exhaustion during high-concurrency mesh operations.

## Frequently Asked Questions

### What is the maximum size of a voice burst packet in BitChat?

Voice burst packets are constrained by `TransportConfig.pttMaxBurstContentBytes` to ensure they fit within single BLE encryption frames. Additionally, each DATA packet carries at most `VoiceBurstPacket.maxFramesPerPacket` (8) AAC frames, limiting the payload size and ensuring real-time latency targets are met during mesh transmission.

### How does BitChat prevent packet collisions from different senders?

The assembly system uses a composite `AssemblyKey` consisting of **peerID**, **scope**, and **burstID**. Since the peerID is cryptographically bound to the Noise session or signed message context, a START packet from a different sender cannot hijack an existing assembly—the key mismatch routes it to a new assembly instance. This isolation prevents cross-traffic corruption in busy mesh networks.

### What audio codec does BitChat use for voice bursts?

The codec is negotiated in the START packet's payload as a UInt8 identifier, with `.aacLC16kMono` (AAC Low Complexity, 16kHz, mono) being the standard implementation. The receiving side inspects `packet.kind` when decoding START to initialize the appropriate decoder before processing subsequent DATA packets containing raw AAC frame payloads.

### How are incomplete voice bursts handled?

Transmission aborts are signaled via CANCELED packets (flag `0x04`). When `ChatLiveVoiceCoordinator` processes a CANCELED packet, it immediately discards the associated assembly from its dictionary and deletes any partially written temporary audio files. This cleanup mechanism prevents storage pollution from interrupted push-to-talk sessions where the sender released the button prematurely or lost connectivity.