How the "Every Programmer Should Know" Repository Explains Latency in Computing

The mtdvio/every-programmer-should-know repository explains latency in computing through a curated, resource-driven approach in its README.md, linking to an interactive infographic covering perceptual UI thresholds and a comprehensive list of hardware latency benchmarks.

Latency in computing represents one of the most critical yet often misunderstood concepts in software engineering, affecting everything from user interface responsiveness to distributed system architecture. Rather than providing a lengthy treatise, the open-source repository mtdvio/every-programmer-should-know takes a curated approach to explaining latency in computing. In the Latency section of the project's README.md, maintainers provide precisely two high-impact external references that together cover both human perception and hardware reality.

Where Latency in Computing Lives in the Repository

All latency-related content in the mtdvio/every-programmer-should-know repository resides in a single location. The project maintains its curated list of essential programming knowledge in the root README.md file, with latency in computing receiving dedicated treatment under its own heading.

The Curated Resources on Latency

The repository's README.md (direct section link) lists two specific resources that explain latency in computing from complementary angles:

  1. Interactive Latency Infographics – A visual guide from UC Berkeley that maps perceptual thresholds in human-computer interaction, illustrating that 100 ms feels instantaneous while delays exceeding 1 second create noticeable interruptions.

  2. Latency Numbers Every Programmer Should Know – A community-maintained gist cataloging concrete latency figures for hardware operations, from L1 cache hits (0.5 ns) to cross-country network round-trips (150 ms).

Understanding Latency Through the Repository's References

The repository's approach to explaining latency in computing emphasizes practical application over theoretical abstraction. By pairing perceptual psychology with hardware metrics, the curated links enable developers to make informed decisions about both user interface design and system architecture.

Perceptual Latency and User Experience

The Berkeley interactive infographics address the human side of latency in computing. According to this resource, perceptual thresholds determine whether users perceive a system as responsive:

  • 100 milliseconds – The threshold below which actions feel instantaneous to users
  • 1 second – The point at which users lose the feeling of direct manipulation and notice a delay
  • 10 seconds – The limit of user attention; beyond this, users context-switch away from the task

Understanding these perceptual boundaries helps developers set appropriate performance budgets for interactive features.

Hardware and Network Latency Benchmarks

The "Latency Numbers Every Programmer Should Know" gist provides the quantitative foundation for reasoning about latency in computing at the system level. The reference lists typical operation latencies that serve as rules of thumb for capacity planning:

  • L1 cache reference: 0.5 ns
  • L2 cache reference: 7 ns
  • Main memory access: 100 ns
  • SSD read: 150 µs
  • Round trip within same datacenter: 500 µs
  • Send packet CA→Netherlands→CA: 150 ms

These figures enable developers to estimate end-to-end latency for request pipelines and identify bottlenecks before they implement solutions.

Practical Applications of Latency Knowledge

The repository's curated approach to latency in computing enables two primary workflows: analytical estimation before writing code and empirical validation after implementation.

Estimating Pipeline Latency

Using the hardware benchmarks from the gist, developers can model the expected latency of a request before building the system. The following Python example demonstrates how to apply these latency in computing benchmarks to estimate total pipeline latency:


# Example: Estimate total latency of a request pipeline (Python)

# Numbers sourced from "Latency Numbers Every Programmer Should Know"

#   L1 cache hit:      0.5 ns

#   L2 cache hit:      7   ns

#   RAM access:       100   ns

#   SSD read:         150 µs

#   1 Gbps network:   8   ms per MB

def estimate_latency(payload_mb: float) -> float:
    """Return an approximate end-to-end latency in milliseconds."""
    l1 = 0.5e-6          # ns → ms

    l2 = 7e-6
    ram = 100e-6
    ssd = 150e-3
    network_per_mb = 8.0  # ms per MB at 1 Gbps

    # Assume one cache hit each, one RAM read, one SSD read,

    # and network transmission of the payload size.

    total_ms = (l1 + l2 + ram + ssd + payload_mb * network_per_mb)
    return total_ms

print(f"Estimated latency for 2 MB payload: {estimate_latency(2):.2f} ms")

This analytical approach helps teams set realistic service-level objectives (SLOs) before committing to specific architectures.

Measuring Real-World Latency

After implementation, developers should validate their estimates against actual measurements. The following Go example demonstrates how to measure HTTP request latency and compare it against the expected network latency benchmarks:

// Example: Measuring real-world latency of a simple HTTP request (Go)
package main

import (
    "net/http"
    "time"
    "log"
)

func main() {
    start := time.Now()
    resp, err := http.Get("https://example.com")
    if err != nil {
        log.Fatalf("request failed: %v", err)
    }
    defer resp.Body.Close()
    elapsed := time.Since(start)

    // Compare against the "network latency" numbers from the gist
    // (~40 ms round-trip for a typical internet request)
    log.Printf("Actual request latency: %v (≈ %.2f ms)", elapsed, float64(elapsed.Nanoseconds())/1e6)
}

By comparing empirical measurements against the theoretical latency in computing benchmarks, teams can identify when their systems deviate from expected performance and optimize accordingly.

Summary

The mtdvio/every-programmer-should-know repository explains latency in computing through a curated, resource-driven approach rather than lengthy exposition. Key takeaways include:

  • The repository's README.md contains a dedicated Latency section linking to two essential external resources.
  • Interactive Latency Infographics from UC Berkeley provide perceptual thresholds for UI design, establishing that 100 ms feels instantaneous while delays over 1 second disrupt user flow.
  • Latency Numbers Every Programmer Should Know offers concrete hardware benchmarks—from 0.5 ns L1 cache hits to 150 ms cross-continental network hops—enabling quantitative system design.
  • Developers can apply these latency in computing principles through analytical estimation (modeling pipelines before coding) and empirical validation (benchmarking implementations against expected values).

Frequently Asked Questions

What is latency in computing?

Latency in computing refers to the time delay between a cause and its effect in a digital system, typically measured as the round-trip time for a data packet, instruction, or user action to complete its cycle. In the context of the every-programmer-should-know repository, latency encompasses both perceptual delays that affect user experience and hardware-level delays that determine system throughput.

The repository maintains a curated list of high-quality external resources rather than duplicating content because latency in computing is a well-documented field with established authoritative sources. By linking to the UC Berkeley interactive infographics and the community-maintained latency numbers gist, the repository ensures developers access continuously updated, peer-reviewed data without requiring the maintainers to update hardware benchmarks as technology evolves.

How can I use the latency numbers for system design?

You can use the latency numbers from the repository's recommended gist to build quantitative performance models before writing code, identifying potential bottlenecks in your proposed architecture. For example, if your design requires reading from SSD (150 µs) then transmitting over a cross-country network (150 ms), you can calculate that network latency dominates by three orders of magnitude and optimize by reducing round-trips or moving data closer to users.

What is the difference between perceptual latency and hardware latency?

Perceptual latency refers to the delays that humans consciously notice when interacting with software, such as the 100 ms threshold where actions feel instantaneous versus the 1 second threshold where users lose their sense of direct manipulation. Hardware latency, conversely, measures the physical time required for electronic operations—such as CPU cache accesses, memory reads, or network packet transmission—regardless of human perception, typically measured in nanoseconds to milliseconds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →