How Does Kafka Achieve High Throughput and Low Latency?
Kafka achieves high throughput and low latency through sequential I/O operations, zero-copy networking via the sendfile() system call, immutable append-only logs, and partition-level parallelism that eliminates disk seek penalties and minimizes CPU overhead.
The ByteByteGoHq/system-design-101 repository documents how Kafka processes millions of messages per second while maintaining millisecond-level latency. Unlike traditional message queues that rely on random disk access and multiple data copies between kernel and user space, Kafka treats logs as the core abstraction, leveraging hardware-optimized sequential writes and operating system primitives to maximize performance.
Sequential I/O and Zero-Copy Architecture
According to data/guides/why-is-kafka-fast.md, Kafka's performance foundation rests on two OS-level optimizations: sequential disk access and zero-copy transfer.
Sequential Disk Writes
Kafka writes data to disk in large, contiguous batches rather than random small writes. This design leverages the natural speed of modern storage devices—both spinning disks and SSDs—for sequential operations, drastically reducing seek time and allowing the OS page cache to stream data efficiently.
By avoiding random writes, the broker can accept far more produce requests per second. Consumers can begin reading as soon as data lands on the log without waiting for expensive random-access operations.
Zero-Copy Transfer with sendfile()
Kafka uses the sendfile() system call to move data from the OS page cache directly to the network socket, bypassing user-space copies entirely. The data path becomes disk → OS cache → network card instead of the traditional disk → OS cache → Kafka process → network buffer → network card.
This zero-copy mechanism eliminates multiple memory copies per message and reduces CPU context switches, freeing CPU cycles for handling additional requests. The reduced copy overhead translates to sub-millisecond delivery from broker to consumer.
Append-Only Log Structure
Each Kafka partition is an immutable, append-only file. Writes never modify existing data, so they never contend with reads.
This append-only design means writers never block readers; the broker can keep appending at the speed of the underlying storage. Consumers can read from any offset without lock contention, leading to predictable, low-latency reads regardless of write volume.
Batching, Compression, and Parallel Processing
As detailed in data/guides/how-do-message-queue-architectures-evolve.md and data/guides/the-ultimate-kafka-101-you-cannot-miss.md, Kafka scales throughput through batching and horizontal partitioning.
Producer Batching and Compression
Producers batch many records into a single request and optionally compress them using algorithms like LZ4 or Snappy. The broker writes the batch as one unit.
Fewer network round-trips and fewer disk writes per message increase overall throughput. A single batch arrives at the consumer in one go, reducing per-message latency overhead.
Partition-Level Parallelism
A topic can be split into many partitions, each handled by a separate broker thread or even separate broker instances. This partition parallelism spreads write and read workloads across many CPU cores and disks, scaling linearly with the number of partitions.
Consumer groups can read from multiple partitions in parallel, cutting end-to-end latency by distributing the workload across many consumers simultaneously.
Configurable Durability for Latency Tuning
The acks setting lets producers decide whether they need confirmation from all replicas or just the leader. When acks=0 or acks=1, the producer can proceed without waiting for costly quorum writes, boosting raw throughput. For latency-critical pipelines, producers can trade durability for faster acknowledgement, while acks=all provides stronger guarantees when throughput can be sacrificed.
Optimizing Kafka Clients for Throughput and Latency
The following Java and Python examples demonstrate the configuration parameters that directly influence Kafka's performance characteristics.
Java Producer Configuration
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("key.serializer", "org.apache.kafka.common.serialization.StringSerializer");
props.put("value.serializer", "org.apache.kafka.common.serialization.StringSerializer");
// High‑throughput: enable batching and compression
props.put("batch.size", 32768); // 32 KB batches
props.put("linger.ms", 5); // Wait up to 5 ms for a batch
props.put("compression.type", "lz4");
// Low‑latency: ack after leader only
props.put("acks", "1");
KafkaProducer<String, String> producer = new KafkaProducer<>(props);
for (int i = 0; i < 1_000_000; i++) {
producer.send(new ProducerRecord<>("orders", Integer.toString(i), "payload"));
}
producer.flush();
producer.close();
Java Consumer Configuration
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("group.id", "order‑service");
props.put("key.deserializer", "org.apache.kafka.common.serialization.StringDeserializer");
props.put("value.deserializer","org.apache.kafka.common.serialization.StringDeserializer");
// Low‑latency: poll as fast as possible
props.put("max.poll.records", 100);
props.put("fetch.min.bytes", 1);
props.put("fetch.max.wait.ms", 50);
KafkaConsumer<String, String> consumer = new KafkaConsumer<>(props);
consumer.subscribe(List.of("orders"));
while (true) {
ConsumerRecords<String, String> records = consumer.poll(Duration.ofMillis(10));
for (ConsumerRecord<String, String> r : records) {
process(r.value()); // fast, in‑memory handling
}
consumer.commitAsync();
}
Python Producer Configuration
from confluent_kafka import Producer
conf = {
'bootstrap.servers': 'broker1:9092,broker2:9092',
'batch.num.messages': 1000,
'linger.ms': 5,
'compression.type': 'lz4',
'acks': '1'
}
producer = Producer(conf)
for i in range(1_000_000):
producer.produce('orders', key=str(i), value='payload')
producer.flush()
Python Consumer Configuration
from confluent_kafka import Consumer
conf = {
'bootstrap.servers': 'broker1:9092,broker2:9092',
'group.id': 'order-service',
'auto.offset.reset': 'earliest',
'max.poll.interval.ms': 100,
'fetch.min.bytes': 1,
'fetch.max.wait.ms': 50
}
consumer = Consumer(conf)
consumer.subscribe(['orders'])
while True:
msgs = consumer.poll(0.01)
if msgs is None: continue
for m in msgs:
process(m.value())
consumer.commit(asynchronous=True)
Key configuration insights:
batch.sizeandlinger.mscontrol how many records are grouped before a disk write—larger batches increase throughput.compression.typereduces network I/O at the cost of CPU; LZ4 offers a favorable latency-throughput trade-off.acks=1(leader-only) provides low latency; set toallfor stronger durability when throughput can be sacrificed.
Summary
- Sequential I/O and zero-copy transfer (
sendfile()) eliminate random disk seeks and user-space memory copies, allowing Kafka to process millions of messages per second with sub-millisecond latency. - Append-only logs ensure writes never block reads, providing predictable low-latency access regardless of write volume.
- Batching, compression, and partition parallelism distribute workload across CPU cores and disks, enabling linear scalability of throughput.
- Configurable durability via the
acksparameter allows producers to trade between strict consistency and low-latency delivery. - Practical client configuration of
batch.size,linger.ms, andcompression.typedirectly controls the throughput-latency trade-off in production workloads.
Frequently Asked Questions
What makes Kafka faster than traditional message queues?
Traditional message queues often rely on random disk I/O and store messages in memory or databases that require multiple data copies between kernel and user space. Kafka achieves superior performance by treating storage as an append-only log with sequential I/O and implementing zero-copy networking via sendfile(), which allows the OS to transfer data directly from page cache to network socket without copying through the Kafka broker process.
How does Kafka's zero-copy mechanism work?
Kafka uses the Linux sendfile() system call to establish a direct data path from disk to network. When a consumer requests data, the broker instructs the OS to transfer bytes from the page cache (where the log file resides) directly to the network card buffer, bypassing the traditional route through user-space buffers. This zero-copy approach eliminates multiple memory copies and CPU context switches, reducing latency to sub-millisecond levels while freeing CPU cycles for additional requests.
Can Kafka guarantee both high throughput and low latency simultaneously?
Kafka can achieve both high throughput and low latency, but these goals require different configuration trade-offs. High throughput benefits from large batch.size, longer linger.ms, and compression, which increase latency slightly to maximize resource efficiency. Low latency requires smaller batches, acks=1 (leader-only acknowledgment), and minimal fetch.max.wait.ms. For optimal results, architects often deploy separate topic configurations or even separate clusters for latency-sensitive versus throughput-heavy workloads.
What configuration changes reduce Kafka latency the most?
The most impactful latency reductions come from adjusting producer and consumer acknowledgment settings and fetch behavior. Set acks=1 on the producer to avoid waiting for replica synchronization, configure fetch.min.bytes=1 and fetch.max.wait.ms to low values (e.g., 50ms) on consumers to ensure immediate delivery of available records, and reduce linger.ms to near zero to minimize batching delays. Additionally, disabling compression or using LZ4 (which balances CPU and compression ratio) reduces processing overhead compared to slower algorithms like GZIP.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →