# How Does Kafka Achieve High Throughput and Low Latency?

> Discover how Kafka achieves high throughput and low latency with sequential I/O, zero-copy networking, immutable logs, and partition parallelism. Optimize your streaming data.

- Repository: [ByteByteGoHq/system-design-101](https://github.com/ByteByteGoHq/system-design-101)
- Tags: deep-dive
- Published: 2026-02-28

---

**Kafka achieves high throughput and low latency through sequential I/O operations, zero-copy networking via the `sendfile()` system call, immutable append-only logs, and partition-level parallelism that eliminates disk seek penalties and minimizes CPU overhead.**

The ByteByteGoHq/system-design-101 repository documents how Kafka processes millions of messages per second while maintaining millisecond-level latency. Unlike traditional message queues that rely on random disk access and multiple data copies between kernel and user space, Kafka treats logs as the core abstraction, leveraging hardware-optimized sequential writes and operating system primitives to maximize performance.

## Sequential I/O and Zero-Copy Architecture

According to [`data/guides/why-is-kafka-fast.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/why-is-kafka-fast.md), Kafka's performance foundation rests on two OS-level optimizations: sequential disk access and zero-copy transfer.

### Sequential Disk Writes

Kafka writes data to disk in large, contiguous batches rather than random small writes. This design leverages the natural speed of modern storage devices—both spinning disks and SSDs—for sequential operations, drastically reducing seek time and allowing the OS page cache to stream data efficiently.

By avoiding random writes, the broker can accept far more produce requests per second. Consumers can begin reading as soon as data lands on the log without waiting for expensive random-access operations.

### Zero-Copy Transfer with sendfile()

Kafka uses the `sendfile()` system call to move data from the OS page cache directly to the network socket, bypassing user-space copies entirely. The data path becomes **disk → OS cache → network card** instead of the traditional **disk → OS cache → Kafka process → network buffer → network card**.

This zero-copy mechanism eliminates multiple memory copies per message and reduces CPU context switches, freeing CPU cycles for handling additional requests. The reduced copy overhead translates to sub-millisecond delivery from broker to consumer.

## Append-Only Log Structure

Each Kafka partition is an immutable, append-only file. Writes never modify existing data, so they never contend with reads.

This append-only design means writers never block readers; the broker can keep appending at the speed of the underlying storage. Consumers can read from any offset without lock contention, leading to predictable, low-latency reads regardless of write volume.

## Batching, Compression, and Parallel Processing

As detailed in [`data/guides/how-do-message-queue-architectures-evolve.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/how-do-message-queue-architectures-evolve.md) and [`data/guides/the-ultimate-kafka-101-you-cannot-miss.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/the-ultimate-kafka-101-you-cannot-miss.md), Kafka scales throughput through batching and horizontal partitioning.

### Producer Batching and Compression

Producers batch many records into a single request and optionally compress them using algorithms like **LZ4** or **Snappy**. The broker writes the batch as one unit.

Fewer network round-trips and fewer disk writes per message increase overall throughput. A single batch arrives at the consumer in one go, reducing per-message latency overhead.

### Partition-Level Parallelism

A topic can be split into many partitions, each handled by a separate broker thread or even separate broker instances. This partition parallelism spreads write and read workloads across many CPU cores and disks, scaling linearly with the number of partitions.

Consumer groups can read from multiple partitions in parallel, cutting end-to-end latency by distributing the workload across many consumers simultaneously.

## Configurable Durability for Latency Tuning

The **`acks`** setting lets producers decide whether they need confirmation from all replicas or just the leader. When `acks=0` or `acks=1`, the producer can proceed without waiting for costly quorum writes, boosting raw throughput. For latency-critical pipelines, producers can trade durability for faster acknowledgement, while `acks=all` provides stronger guarantees when throughput can be sacrificed.

## Optimizing Kafka Clients for Throughput and Latency

The following Java and Python examples demonstrate the configuration parameters that directly influence Kafka's performance characteristics.

### Java Producer Configuration

```java
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("key.serializer",   "org.apache.kafka.common.serialization.StringSerializer");
props.put("value.serializer", "org.apache.kafka.common.serialization.StringSerializer");

// High‑throughput: enable batching and compression
props.put("batch.size", 32768);            // 32 KB batches
props.put("linger.ms", 5);                // Wait up to 5 ms for a batch
props.put("compression.type", "lz4");

// Low‑latency: ack after leader only
props.put("acks", "1");

KafkaProducer<String, String> producer = new KafkaProducer<>(props);
for (int i = 0; i < 1_000_000; i++) {
    producer.send(new ProducerRecord<>("orders", Integer.toString(i), "payload"));
}
producer.flush();
producer.close();

```

### Java Consumer Configuration

```java
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("group.id", "order‑service");
props.put("key.deserializer",  "org.apache.kafka.common.serialization.StringDeserializer");
props.put("value.deserializer","org.apache.kafka.common.serialization.StringDeserializer");

// Low‑latency: poll as fast as possible
props.put("max.poll.records", 100);
props.put("fetch.min.bytes", 1);
props.put("fetch.max.wait.ms", 50);

KafkaConsumer<String, String> consumer = new KafkaConsumer<>(props);
consumer.subscribe(List.of("orders"));
while (true) {
    ConsumerRecords<String, String> records = consumer.poll(Duration.ofMillis(10));
    for (ConsumerRecord<String, String> r : records) {
        process(r.value());   // fast, in‑memory handling
    }
    consumer.commitAsync();
}

```

### Python Producer Configuration

```python
from confluent_kafka import Producer

conf = {
    'bootstrap.servers': 'broker1:9092,broker2:9092',
    'batch.num.messages': 1000,
    'linger.ms': 5,
    'compression.type': 'lz4',
    'acks': '1'
}
producer = Producer(conf)

for i in range(1_000_000):
    producer.produce('orders', key=str(i), value='payload')
producer.flush()

```

### Python Consumer Configuration

```python
from confluent_kafka import Consumer

conf = {
    'bootstrap.servers': 'broker1:9092,broker2:9092',
    'group.id': 'order-service',
    'auto.offset.reset': 'earliest',
    'max.poll.interval.ms': 100,
    'fetch.min.bytes': 1,
    'fetch.max.wait.ms': 50
}
consumer = Consumer(conf)
consumer.subscribe(['orders'])

while True:
    msgs = consumer.poll(0.01)
    if msgs is None: continue
    for m in msgs:
        process(m.value())
    consumer.commit(asynchronous=True)

```

**Key configuration insights:**

- **`batch.size`** and **`linger.ms`** control how many records are grouped before a disk write—larger batches increase throughput.
- **`compression.type`** reduces network I/O at the cost of CPU; **LZ4** offers a favorable latency-throughput trade-off.
- **`acks=1`** (leader-only) provides low latency; set to `all` for stronger durability when throughput can be sacrificed.

## Summary

- **Sequential I/O** and **zero-copy transfer** (`sendfile()`) eliminate random disk seeks and user-space memory copies, allowing Kafka to process millions of messages per second with sub-millisecond latency.
- **Append-only logs** ensure writes never block reads, providing predictable low-latency access regardless of write volume.
- **Batching, compression, and partition parallelism** distribute workload across CPU cores and disks, enabling linear scalability of throughput.
- **Configurable durability** via the `acks` parameter allows producers to trade between strict consistency and low-latency delivery.
- Practical client configuration of `batch.size`, `linger.ms`, and `compression.type` directly controls the throughput-latency trade-off in production workloads.

## Frequently Asked Questions

### What makes Kafka faster than traditional message queues?

Traditional message queues often rely on random disk I/O and store messages in memory or databases that require multiple data copies between kernel and user space. Kafka achieves superior performance by treating storage as an append-only log with **sequential I/O** and implementing **zero-copy** networking via `sendfile()`, which allows the OS to transfer data directly from page cache to network socket without copying through the Kafka broker process.

### How does Kafka's zero-copy mechanism work?

Kafka uses the Linux `sendfile()` system call to establish a direct data path from disk to network. When a consumer requests data, the broker instructs the OS to transfer bytes from the **page cache** (where the log file resides) directly to the **network card buffer**, bypassing the traditional route through user-space buffers. This zero-copy approach eliminates multiple memory copies and CPU context switches, reducing latency to sub-millisecond levels while freeing CPU cycles for additional requests.

### Can Kafka guarantee both high throughput and low latency simultaneously?

Kafka can achieve both high throughput and low latency, but these goals require different configuration trade-offs. **High throughput** benefits from large `batch.size`, longer `linger.ms`, and compression, which increase latency slightly to maximize resource efficiency. **Low latency** requires smaller batches, `acks=1` (leader-only acknowledgment), and minimal `fetch.max.wait.ms`. For optimal results, architects often deploy separate topic configurations or even separate clusters for latency-sensitive versus throughput-heavy workloads.

### What configuration changes reduce Kafka latency the most?

The most impactful latency reductions come from adjusting producer and consumer acknowledgment settings and fetch behavior. Set **`acks=1`** on the producer to avoid waiting for replica synchronization, configure **`fetch.min.bytes=1`** and **`fetch.max.wait.ms`** to low values (e.g., 50ms) on consumers to ensure immediate delivery of available records, and reduce **`linger.ms`** to near zero to minimize batching delays. Additionally, disabling compression or using **LZ4** (which balances CPU and compression ratio) reduces processing overhead compared to slower algorithms like GZIP.