# Performance Considerations for iroh: Achieving High-Throughput, Low-Latency P2P Transfers

> Discover iroh performance considerations for gigabit throughput and sub-10ms latency. Learn how QUIC multiplexing, Tokio runtimes, zero-copy, and NAT traversal optimize P2P transfers.

- Repository: [number zero/iroh](https://github.com/n0-computer/iroh)
- Tags: performance
- Published: 2026-07-12

---

**Iroh delivers gigabit-class throughput and sub-10ms latency by combining QUIC multiplexing, configurable Tokio runtimes, zero-copy blob handling, and intelligent NAT traversal that prioritizes direct hole-punched connections over relay fallback.**

The `iroh` library from n0-computer/iroh is a Rust-based peer-to-peer networking stack built on top of QUIC. Understanding the performance considerations for iroh requires examining its core architectural decisions, from transport-layer optimizations to memory-efficient data streaming. This guide analyzes the source code to reveal how the library achieves scalable concurrency while maintaining minimal resource overhead.

## Core Performance Architecture

### QUIC Transport and Stream Multiplexing

At the foundation of iroh's performance lies its use of the `quinn` crate for QUIC transport. Unlike TCP, QUIC provides **native stream multiplexing** within a single connection, eliminating head-of-line blocking between independent data flows. In [`iroh/src/transport/quic.rs`](https://github.com/n0-computer/iroh/blob/main/iroh/src/transport/quic.rs), the transport layer leverages QUIC's built-in congestion control and TLS encryption to saturate network links while maintaining security. The implementation supports many concurrent streams over one connection, allowing the library to fully utilize available bandwidth without establishing multiple expensive connections.

### Connection Establishment and Path Selection

Iroh optimizes latency through aggressive **NAT hole-punching** with relay fallback. The `Endpoint` implementation in [`iroh/src/endpoint.rs`](https://github.com/n0-computer/iroh/blob/main/iroh/src/endpoint.rs) attempts direct peer-to-peer connectivity first, falling back to encrypted relay servers only when necessary. Direct paths typically achieve sub-10ms time-to-first-byte (TTFB) on local networks, while relay connections introduce modest additional latency. This dual-path approach ensures that performance degrades gracefully when direct connectivity is impossible, without sacrificing end-to-end encryption.

### Async Runtime Configuration

Each endpoint runs on a dedicated **Tokio runtime** where the number of worker threads directly impacts throughput. The `--workers-per-ep` parameter configures how many threads process I/O and protocol logic. Benchmarks in the codebase demonstrate that allocating two workers per endpoint provides optimal balance between ACK processing and data transmission, though this can be scaled for high-concurrency scenarios.

## Memory Efficiency and Data Handling

### Zero-Copy Blob Transfers

For large data transfers, iroh utilizes **BLAKE3-based content addressing** via the `iroh-blobs` protocol. This design enables zero-copy streaming where data moves directly from network buffers to storage without intermediate copying. As implemented in the blob transfer handlers, this approach minimizes CPU overhead and memory pressure, making multi-terabyte transfers feasible without proportional memory consumption.

### Protocol Composition Overhead

Iroh's architecture allows multiple protocol handlers—such as gossip and blobs—to share a single QUIC connection. This **multi-protocol composition** eliminates the need for separate connections per protocol, reducing connection establishment overhead and allowing better congestion control across the entire application stack.

## Benchmarking with iroh-bench

### Measuring Real-World Performance

The repository includes `iroh-bench`, a comprehensive benchmarking tool located in [`iroh/bench/src/lib.rs`](https://github.com/n0-computer/iroh/blob/main/iroh/bench/src/lib.rs). This utility drives the library under configurable loads to measure actual throughput, latency, and resource utilization. The `client_handler` logic in this file coordinates multiple simultaneous endpoints to stress-test the transport layer.

Key configuration parameters include:

- `--clients`: Number of simultaneous endpoints
- `--streams`: Concurrent streams per client
- `--max-streams`: Upper bound for simultaneous streams
- `--download-size` and `--upload-size`: Payload sizes with SI prefixes
- `--workers-per-ep`: Tokio worker threads per endpoint

### Interpreting Metrics

The statistics collection system in [`iroh/bench/src/stats.rs`](https://github.com/n0-computer/iroh/blob/main/iroh/bench/src/stats.rs) captures granular performance data. The `Stats` struct aggregates per-stream histograms for duration, throughput, TTFB, and chunk processing time.

According to the implementation in [`stats.rs`](https://github.com/n0-computer/iroh/blob/main/stats.rs), **throughput** is calculated as `size / duration` in the `throughput_bps` calculation (lines 33-34), while **TTFB** (Time to First Byte) is captured in `TransferResult::new` (lines 19-20) as the duration until the first byte arrives. Typical benchmark output reveals:

```text
Overall upload stats:
Transferred 4.09 GiB on 8 streams in 5.12 s (0.80 MiB/s)
Time to first byte (TTFB): 12.4 ms
Total chunks: 12345
Average chunk time: 0.41 ms

```

### Programmatic Benchmarking

You can embed benchmark logic directly in your application using the `iroh::bench` modules:

```rust
// Example: Run a simple benchmark from Rust code
use iroh::bench::{self, Opt, Commands};
use clap::Parser;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Configure a benchmark: 4 clients, 4 streams each, 100 MiB download per stream
    let opt = Opt {
        clients: 4,
        streams: 4,
        max_streams: 4,
        download_size: 100 << 20, // 100 MiB
        upload_size: 0,
        ..Default::default()
    };
    // Run the Iroh benchmark directly (same logic used by `iroh-bench`)
    bench::iroh::run(opt).await?;
    Ok(())
}

```

Access detailed statistics after transfers:

```rust
// Example: Access per-stream statistics after a transfer
use iroh::bench::stats::{Stats, TransferResult};

fn report(stats: &Stats) {
    println!("Total bytes transferred: {}", stats.total_size);
    println!("Average throughput: {:.2} MiB/s",
        stats.stream_stats.throughput_hist.mean() as f64 / 1024.0 / 1024.0);
    println!("Mean TTFB: {:.2} ms",
        stats.stream_stats.ttfb_hist.mean() / 1_000.0);
}

```

## Tuning Strategies for Production

### Optimizing Worker Thread Allocation

While single-threaded runtimes suffice for low-throughput scenarios, production deployments handling gigabit traffic should configure `--workers-per-ep 2` or higher. This distributes ACK processing and stream handling across CPU cores, improving congestion-control responsiveness and preventing the runtime from becoming a bottleneck during burst transfers.

### MTU and Chunk Size Considerations

The per-chunk timings recorded in `chunk_time` histograms reveal how Maximum Transmission Unit (MTU) settings affect latency. Smaller MTU values increase chunk counts but maintain low per-chunk latency, beneficial for interactive applications. Larger MTU values reduce protocol overhead and maximize raw throughput, ideal for bulk data transfers. The `--initial-mtu` parameter allows tuning this trade-off for specific network conditions.

## Summary

- **QUIC multiplexing** via `quinn` eliminates head-of-line blocking and enables efficient stream concurrency over single connections.
- **Zero-copy blob handling** using BLAKE3 content addressing minimizes CPU and memory overhead for large transfers.
- **Configurable Tokio workers** (`--workers-per-ep`) allow scaling across CPU cores to match network throughput.
- **Intelligent path selection** prioritizes direct hole-punched connections for minimal latency, with encrypted relay fallback.
- **Built-in instrumentation** via `iroh-bench` and [`stats.rs`](https://github.com/n0-computer/iroh/blob/main/stats.rs) provides granular visibility into TTFB, throughput, and chunk-level performance.

## Frequently Asked Questions

### How does iroh achieve low latency in peer-to-peer connections?

Iroh minimizes latency by attempting NAT hole-punching first to establish direct paths between peers, which typically achieves sub-10ms TTFB on local networks. When direct connectivity fails, it falls back to encrypted relay servers, maintaining security with only modest latency increases. The QUIC transport further reduces latency through 0-RTT connection resumption and improved congestion control compared to TCP.

### What is the recommended configuration for high-throughput transfers?

For high-throughput scenarios, configure `--workers-per-ep 2` to distribute workload across CPU cores, and utilize the zero-copy blob handling in `iroh-blobs` for large data transfers. Ensure your application leverages QUIC's stream multiplexing by opening multiple concurrent streams rather than sequential transfers, allowing the congestion controller to saturate the available bandwidth.

### How does iroh handle memory usage during large file transfers?

Iroh implements zero-copy streaming where data moves directly from network buffers to storage without intermediate copies, significantly reducing memory pressure. The BLAKE3-based content addressing in `iroh-blobs` allows streaming verification of data integrity without loading entire files into memory, enabling efficient handling of multi-terabyte transfers.

### Can I programmatically access performance metrics from iroh?

Yes, the `iroh::bench::stats` module provides programmatic access to detailed performance metrics. You can instantiate `Stats` and `TransferResult` structures to capture histogram data for throughput, TTFB, and chunk processing times, as implemented in [`iroh/bench/src/stats.rs`](https://github.com/n0-computer/iroh/blob/main/iroh/bench/src/stats.rs). This allows integration of performance monitoring directly into your application logic.