# How the Push-to-Talk Voice Feature (VoiceBurstPacket) Functions in the Bitchat Mesh

> Discover how the push-to-talk voice feature VoiceBurstPacket fragments audio into BLE packets and streams them through the mesh for real-time playback.

- Repository: [permissionlesstech/bitchat](https://github.com/permissionlesstech/bitchat)
- Tags: deep-dive
- Published: 2026-08-09

---

**The push-to-talk voice feature fragments live AAC audio into BLE-sized VoiceBurstPackets that are streamed through the mesh, assembled per-burst by sequence number, and played back in real-time before being promoted to permanent voice-note files.**

The **permissionlesstech/bitchat** mesh network implements push-to-talk (PTT) voice messaging through a lightweight packetization protocol designed for low-bandwidth Bluetooth Low Energy (BLE) transport. At the core of this system is the `VoiceBurstPacket` type, which manages the fragmentation, sequencing, and reassembly of live voice bursts across decentralized peers.

## Voice Burst Packet Wire Format

Each `VoiceBurstPacket` represents a single BLE-sized fragment of a live voice burst. The binary wire format is structured as:

```

[burstID: 8 bytes][seq: UInt16 BE][flags: UInt8][payload…]

```

- **`burstID`** (8 bytes): Uniquely identifies a burst for a given sender-scope pair.
- **`seq`** (UInt16 Big Endian): Packet sequence number where `0` indicates **START** and values `≥ 1` represent data packets.
- **`flags`** (UInt8): Indicates packet kind—**START** (`0x01`), **END** (`0x02`), **CANCELED** (`0x04`), or data (`0x00`).
- **payload**: Variable content depending on the flag—codec byte for START packets, AAC frame data for data packets, or total packet count and duration for END packets.

The full definition, encoding logic, and decoding validation reside in [`bitchat/Protocols/VoiceBurstPacket.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Protocols/VoiceBurstPacket.swift).

## Encoding and Decoding Voice Frames

The `VoiceBurstPacket` type provides deterministic serialization methods for mesh transmission.

**`VoiceBurstPacket.encode()`** builds the binary payload by appending the appropriate flag and codec or frame data. ** `VoiceBurstPacket.decode(_:)`** validates packet length, extracts the flag byte, and reconstructs the `Kind` enum (`.start`, `.frames`, `.end`, `.canceled`). The decoder rejects unknown flags or malformed frames, preventing garbage data from reaching the audio pipeline.

## Outgoing Packetization

Before transmission, raw AAC frames are collected and fragmented by `VoiceBurstPacketizer` to respect BLE payload budgets.

The packetizer accumulates encoded frames in an internal buffer. When the buffer reaches `TransportConfig.pttMaxBurstContentBytes` or exceeds the maximum per-packet frame count, it flushes the current frames into a complete `VoiceBurstPacket` and returns the encoded `Data`. This ensures no single packet exceeds the mesh transport's maximum transmission unit.

Implementation details for the buffering and flushing logic are found in [`bitchat/Protocols/VoiceBurstPacket.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Protocols/VoiceBurstPacket.swift) (lines 73–104 and 118–124).

## Transport Layer Integration

The mesh transport exposes a lightweight capability for voice streaming via the `MeshVoiceStreaming` protocol defined in [`bitchat/Services/MeshTransportCapabilities.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/Services/MeshTransportCapabilities.swift):

```swift
protocol MeshVoiceStreaming: AnyObject {
    func sendVoiceFrame(_ burstContent: Data, to peerID: PeerID)
    func sendVoiceFrameBroadcast(_ burstContent: Data)
}

```

Applications use this protocol to fire-and-forget each encoded packet either privately through an established Noise session or as a signed broadcast to all reachable peers. The transport layer handles the underlying routing without requiring the voice coordinator to manage peer connections directly.

## Inbound Burst Assembly

Inbound voice handling is orchestrated by `ChatLiveVoiceCoordinator` in [`bitchat/ViewModels/ChatLiveVoiceCoordinator.swift`](https://github.com/permissionlesstech/bitchat/blob/main/bitchat/ViewModels/ChatLiveVoiceCoordinator.swift). When a packet arrives, `handleVoiceFramePayload(from:payload:timestamp:)` decodes the raw data using `VoiceBurstPacket.decode(_:)`.

If the burst is new, `makeAssembly` creates a per-burst state holder that opens a temporary file named `voice_live_<burstID>_… .aac` and registers the initial chat bubble. Data packets (`.frames`) are buffered by sequence number; `drainInOrder` delivers frames in ascending sequence, writes them to the live AAC file, and queues them for playback via `PTTBurstPlayer`.

Upon receiving an **END** packet, the coordinator stores the total packet count and duration. Once all expected data packets are received—or if a timeout or gap-skip occurs—the burst is **finalized**. The `finalize(_:)` method closes the temporary file, renames it to a permanent `voice_<hex>.m4a` name, updates the chat bubble to reference the permanent file, and removes the assembly from active memory.

## Security and Quality of Service

The push-to-talk implementation includes several safeguards against abuse and corruption:

- **Validation**: The decoder rejects packets with unknown flags or truncated payloads before they reach the assembly logic.
- **Rate limiting**: `TransportConfig.pttInboundMaxBytesPerSecond` and `pttMaxBurstBytes` cap incoming traffic to prevent flood attacks.
- **Codec negotiation**: START packets explicitly declare the codec (e.g., `aacLC16kMono`). If an assembly receives an unsupported codec identifier, the burst is cancelled immediately to prevent decoder errors.

## Lifecycle Management

Voice burst assemblies are ephemeral state machines that require careful cleanup. The coordinator purges assemblies during panic resets, idle timeouts, or when a **CANCELED** flag is received. This ensures no orphaned temporary files remain in the filesystem and memory pressure is minimized during extended mesh sessions.

## Summary

- **VoiceBurstPacket** serializes live voice into BLE-compatible binary frames with an 8-byte burst ID, sequence number, and flags indicating START, END, or CANCELED states.
- **`VoiceBurstPacketizer`** manages outgoing fragmentation, ensuring packets respect `TransportConfig.pttMaxBurstContentBytes`.
- **`MeshVoiceStreaming`** provides the transport abstraction for both unicast and broadcast voice frame delivery.
- **`ChatLiveVoiceCoordinator`** handles inbound reassembly, out-of-order buffering via `drainInOrder`, live playback, and promotion of temporary AAC files to permanent M4A voice notes.
- **Security** is enforced through strict flag validation, codec verification, and byte-rate limiting to prevent mesh flooding.

## Frequently Asked Questions

### What is the maximum payload size for a VoiceBurstPacket?

The maximum payload is determined by `TransportConfig.pttMaxBurstContentBytes`, which defines the BLE payload budget. The `VoiceBurstPacketizer` automatically fragments AAC frames to ensure no packet exceeds this limit, typically configured to fit within standard BLE MTU constraints while accounting for mesh protocol overhead.

### How does Bitchat handle packet loss during voice transmission?

The `ChatLiveVoiceCoordinator` uses sequence numbers to detect gaps. If a packet arrives out of order, it is buffered until the missing sequence arrives or a timeout occurs. The `drainInOrder` method only writes contiguous frames to the live AAC file, ensuring the audio decoder receives valid sequential data. Missing frames result in brief audio gaps rather than decoder corruption.

### What audio codec does the push-to-talk feature use?

The system currently uses **AAC-LC 16kHz mono** (`aacLC16kMono`), declared in the START packet's codec field. The receiver validates this codec identifier during assembly initialization; unsupported codecs trigger immediate cancellation of the burst to prevent playback errors.

### How are voice bursts secured in the mesh network?

Voice frames transmit through established Noise protocol sessions for private delivery, or as cryptographically signed broadcasts. The `MeshVoiceStreaming` protocol abstracts the encryption layer, ensuring that `VoiceBurstPacket` contents are opaque to intermediary nodes. Additionally, rate limiting and burst size caps prevent malicious peers from exhausting bandwidth with malformed or oversized voice streams.