How the Push-to-Talk Voice Feature (VoiceBurstPacket) Functions in the Bitchat Mesh

The push-to-talk voice feature fragments live AAC audio into BLE-sized VoiceBurstPackets that are streamed through the mesh, assembled per-burst by sequence number, and played back in real-time before being promoted to permanent voice-note files.

The permissionlesstech/bitchat mesh network implements push-to-talk (PTT) voice messaging through a lightweight packetization protocol designed for low-bandwidth Bluetooth Low Energy (BLE) transport. At the core of this system is the VoiceBurstPacket type, which manages the fragmentation, sequencing, and reassembly of live voice bursts across decentralized peers.

Voice Burst Packet Wire Format

Each VoiceBurstPacket represents a single BLE-sized fragment of a live voice burst. The binary wire format is structured as:


[burstID: 8 bytes][seq: UInt16 BE][flags: UInt8][payload…]

  • burstID (8 bytes): Uniquely identifies a burst for a given sender-scope pair.
  • seq (UInt16 Big Endian): Packet sequence number where 0 indicates START and values ≥ 1 represent data packets.
  • flags (UInt8): Indicates packet kind—START (0x01), END (0x02), CANCELED (0x04), or data (0x00).
  • payload: Variable content depending on the flag—codec byte for START packets, AAC frame data for data packets, or total packet count and duration for END packets.

The full definition, encoding logic, and decoding validation reside in bitchat/Protocols/VoiceBurstPacket.swift.

Encoding and Decoding Voice Frames

The VoiceBurstPacket type provides deterministic serialization methods for mesh transmission.

VoiceBurstPacket.encode() builds the binary payload by appending the appropriate flag and codec or frame data. ** VoiceBurstPacket.decode(_:)** validates packet length, extracts the flag byte, and reconstructs the Kind enum (.start, .frames, .end, .canceled). The decoder rejects unknown flags or malformed frames, preventing garbage data from reaching the audio pipeline.

Outgoing Packetization

Before transmission, raw AAC frames are collected and fragmented by VoiceBurstPacketizer to respect BLE payload budgets.

The packetizer accumulates encoded frames in an internal buffer. When the buffer reaches TransportConfig.pttMaxBurstContentBytes or exceeds the maximum per-packet frame count, it flushes the current frames into a complete VoiceBurstPacket and returns the encoded Data. This ensures no single packet exceeds the mesh transport's maximum transmission unit.

Implementation details for the buffering and flushing logic are found in bitchat/Protocols/VoiceBurstPacket.swift (lines 73–104 and 118–124).

Transport Layer Integration

The mesh transport exposes a lightweight capability for voice streaming via the MeshVoiceStreaming protocol defined in bitchat/Services/MeshTransportCapabilities.swift:

protocol MeshVoiceStreaming: AnyObject {
    func sendVoiceFrame(_ burstContent: Data, to peerID: PeerID)
    func sendVoiceFrameBroadcast(_ burstContent: Data)
}

Applications use this protocol to fire-and-forget each encoded packet either privately through an established Noise session or as a signed broadcast to all reachable peers. The transport layer handles the underlying routing without requiring the voice coordinator to manage peer connections directly.

Inbound Burst Assembly

Inbound voice handling is orchestrated by ChatLiveVoiceCoordinator in bitchat/ViewModels/ChatLiveVoiceCoordinator.swift. When a packet arrives, handleVoiceFramePayload(from:payload:timestamp:) decodes the raw data using VoiceBurstPacket.decode(_:).

If the burst is new, makeAssembly creates a per-burst state holder that opens a temporary file named voice_live_<burstID>_… .aac and registers the initial chat bubble. Data packets (.frames) are buffered by sequence number; drainInOrder delivers frames in ascending sequence, writes them to the live AAC file, and queues them for playback via PTTBurstPlayer.

Upon receiving an END packet, the coordinator stores the total packet count and duration. Once all expected data packets are received—or if a timeout or gap-skip occurs—the burst is finalized. The finalize(_:) method closes the temporary file, renames it to a permanent voice_<hex>.m4a name, updates the chat bubble to reference the permanent file, and removes the assembly from active memory.

Security and Quality of Service

The push-to-talk implementation includes several safeguards against abuse and corruption:

  • Validation: The decoder rejects packets with unknown flags or truncated payloads before they reach the assembly logic.
  • Rate limiting: TransportConfig.pttInboundMaxBytesPerSecond and pttMaxBurstBytes cap incoming traffic to prevent flood attacks.
  • Codec negotiation: START packets explicitly declare the codec (e.g., aacLC16kMono). If an assembly receives an unsupported codec identifier, the burst is cancelled immediately to prevent decoder errors.

Lifecycle Management

Voice burst assemblies are ephemeral state machines that require careful cleanup. The coordinator purges assemblies during panic resets, idle timeouts, or when a CANCELED flag is received. This ensures no orphaned temporary files remain in the filesystem and memory pressure is minimized during extended mesh sessions.

Summary

  • VoiceBurstPacket serializes live voice into BLE-compatible binary frames with an 8-byte burst ID, sequence number, and flags indicating START, END, or CANCELED states.
  • VoiceBurstPacketizer manages outgoing fragmentation, ensuring packets respect TransportConfig.pttMaxBurstContentBytes.
  • MeshVoiceStreaming provides the transport abstraction for both unicast and broadcast voice frame delivery.
  • ChatLiveVoiceCoordinator handles inbound reassembly, out-of-order buffering via drainInOrder, live playback, and promotion of temporary AAC files to permanent M4A voice notes.
  • Security is enforced through strict flag validation, codec verification, and byte-rate limiting to prevent mesh flooding.

Frequently Asked Questions

What is the maximum payload size for a VoiceBurstPacket?

The maximum payload is determined by TransportConfig.pttMaxBurstContentBytes, which defines the BLE payload budget. The VoiceBurstPacketizer automatically fragments AAC frames to ensure no packet exceeds this limit, typically configured to fit within standard BLE MTU constraints while accounting for mesh protocol overhead.

How does Bitchat handle packet loss during voice transmission?

The ChatLiveVoiceCoordinator uses sequence numbers to detect gaps. If a packet arrives out of order, it is buffered until the missing sequence arrives or a timeout occurs. The drainInOrder method only writes contiguous frames to the live AAC file, ensuring the audio decoder receives valid sequential data. Missing frames result in brief audio gaps rather than decoder corruption.

What audio codec does the push-to-talk feature use?

The system currently uses AAC-LC 16kHz mono (aacLC16kMono), declared in the START packet's codec field. The receiver validates this codec identifier during assembly initialization; unsupported codecs trigger immediate cancellation of the burst to prevent playback errors.

How are voice bursts secured in the mesh network?

Voice frames transmit through established Noise protocol sessions for private delivery, or as cryptographically signed broadcasts. The MeshVoiceStreaming protocol abstracts the encryption layer, ensuring that VoiceBurstPacket contents are opaque to intermediary nodes. Additionally, rate limiting and burst size caps prevent malicious peers from exhausting bandwidth with malformed or oversized voice streams.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →