Common Causes of Deadlocks and How to Prevent Them

Deadlocks occur when two or more processes each hold a lock that the other needs, creating a circular wait that can be prevented by eliminating one of the four Coffman conditions through resource ordering, timeouts, or lock-free algorithms.

A deadlock is a state of permanent blocking in concurrent systems where processes compete for shared resources. According to the ByteByteGoHq/system-design-101 repository, understanding the mechanics behind resource contention is essential for building reliable distributed systems. This guide examines the root causes of deadlocks and provides production-ready prevention strategies based on the Coffman conditions.

What Is a Deadlock?

In data/guides/what-is-a-deadlock.md, a deadlock is defined as a situation where two or more processes (or threads, transactions) are unable to proceed because each is waiting for the other to release a resource. The theoretical foundation for understanding deadlocks rests on the Coffman conditions, which specify that all four conditions must hold simultaneously for a deadlock to occur.

The Four Coffman Conditions

The following table summarizes the four necessary conditions for deadlock as documented in the source repository:

Condition Description
Mutual Exclusion At least one resource can be held by only one process at a time.
Hold‑and‑Wait A process holds one resource while waiting for another.
No Preemption Resources cannot be forcibly taken away from a process.
Circular Wait A circular chain of processes exists where each waits for a resource held by the next.

When all four conditions are present, the system is vulnerable to deadlock.

Common Causes of Deadlocks in Production Systems

Based on the analysis of concurrent programming patterns, four primary causes frequently trigger deadlocks in production environments:

Improper Lock Ordering

When two threads acquire locks in opposite orders, they create a circular wait. For example, Thread 1 locks Resource A then B, while Thread 2 locks B then A. This inversion directly violates proper resource ordering protocols and satisfies the circular wait condition.

Long‑Running Transactions

Transactions that hold locks for extended periods increase the probability that another transaction will request the same resources. The longer the hold‑and‑wait state persists, the higher the deadlock risk, particularly in database systems.

Nested Locks with Hold‑and‑Wait

Code that acquires a lock, performs work, then attempts to acquire another lock without releasing the first creates a nested dependency. This pattern satisfies the hold‑and‑wait condition and is a common source of deadlocks in application code.

Lack of Timeouts

If a process never gives up waiting for a resource, the deadlock persists indefinitely. Systems without bounded waiting mechanisms cannot break the circular wait condition automatically, allowing deadlocks to remain until manual intervention or system restart.

How to Prevent Deadlocks

Preventing deadlocks requires eliminating at least one of the Coffman conditions. The following strategies target the most common conditions—circular wait and hold‑and‑wait—to reduce deadlock likelihood in production systems.

Enforce Global Resource Ordering

Impose a strict total order on all lock types and ensure every thread requests locks following that sequence. This approach eliminates the circular wait condition by construction.

In data/guides/what-is-a-deadlock.md, the following Java example demonstrates strict lock ordering:

// Global lock ordering: first acquire lockA, then lockB
private final ReentrantLock lockA = new ReentrantLock();
private final ReentrantLock lockB = new ReentrantLock();

void safeMethod() {
    lockA.lock();               // Acquire in defined order
    try {
        lockB.lock();           // Always lockB after lockA
        try {
            // critical section using both resources
        } finally {
            lockB.unlock();
        }
    } finally {
        lockA.unlock();
    }
}

If every thread follows the same order, a circular wait cannot form.

Implement Timeouts and Lease‑Based Locks

Use bounded waiting periods to break the hold‑and‑wait condition. If a thread cannot acquire a lock within a specified duration, it releases its current locks and retries, preventing indefinite blocking.

The following Python example from the source repository illustrates timeout usage:

import threading

lock = threading.Lock()

def work():
    # Try to acquire the lock for at most 2 seconds

    acquired = lock.acquire(timeout=2)
    if not acquired:
        print("Could not get lock – aborting to avoid deadlock")
        return
    try:
        # critical section

        pass
    finally:
        lock.release()

The timeout parameter forces the thread to give up, breaking the hold‑and‑wait condition.

Adopt Lock‑Free and Non‑Blocking Algorithms

Eliminate locks entirely by using data structures that rely on atomic primitives and compare‑and‑swap operations. This approach removes the mutual exclusion and circular wait conditions.

As documented in data/guides/blocking-vs-non-blocking-queue.md, non‑blocking designs avoid deadlocks altogether. The following Go example demonstrates a channel‑based approach:

// Use a channel (non‑blocking) instead of a lock‑protected queue
jobs := make(chan Job, 100)

func worker() {
    for job := range jobs {
        process(job) // no explicit lock needed
    }
}

By avoiding explicit locks, the system eliminates the root causes of deadlock.

Deadlock Detection and Recovery

For systems where prevention is impractical, implement periodic deadlock detection algorithms. When a deadlock is detected, the system selects a victim process to terminate or rollback, breaking the circular wait.

Database engines often implement this strategy automatically, monitoring wait‑for graphs to identify cycles and aborting transactions to resolve deadlocks.

Summary

  • Deadlocks require four Coffman conditions: Mutual Exclusion, Hold‑and‑Wait, No Preemption, and Circular Wait must all be present simultaneously for a deadlock to occur.
  • Common causes include improper lock ordering, long‑running transactions, nested locks with hold‑and‑wait patterns, and lack of timeouts.
  • Prevention strategies focus on eliminating Circular Wait through global resource ordering, breaking Hold‑and‑Wait with timeouts, removing locks entirely via non‑blocking algorithms, or implementing detection and recovery mechanisms.
  • Practical implementations in Java, Python, and Go demonstrate how strict ordering, bounded waits, and lock‑free channels eliminate deadlock risks in production code.

Frequently Asked Questions

What are the four Coffman conditions for deadlock?

The four Coffman conditions are Mutual Exclusion (resources cannot be shared), Hold‑and‑Wait (processes hold resources while waiting for others), No Preemption (resources cannot be forcibly released), and Circular Wait (a cycle exists in the resource allocation graph). All four must be true simultaneously for a deadlock to occur, as documented in data/guides/what-is-a-deadlock.md.

How does improper lock ordering cause deadlocks?

Improper lock ordering creates a circular wait condition. When Thread A acquires Lock 1 then Lock 2, while Thread B acquires Lock 2 then Lock 1, each thread holds one lock while waiting for the other. This inversion prevents either thread from proceeding. Enforcing a global order—such as always acquiring Lock 1 before Lock 2—eliminates this circular dependency.

Can deadlocks occur in distributed systems?

Yes, deadlocks can occur in distributed systems when multiple nodes hold resources and wait for resources held by other nodes across the network. Distributed deadlocks are harder to detect because they span multiple machines and network partitions can mask the cyclic wait conditions. Prevention strategies include timeout‑based leases, global ordering of distributed resources, and distributed deadlock detection algorithms that construct global wait‑for graphs.

What is the difference between deadlock prevention and deadlock avoidance?

Deadlock prevention eliminates one of the four Coffman conditions structurally—such as enforcing resource ordering to prevent circular wait or using timeouts to break hold‑and‑wait. Deadlock avoidance, such as the Banker’s Algorithm, does not impose strict structural constraints but instead analyzes resource allocation states dynamically to ensure the system never enters an unsafe state. Avoidance requires knowledge of future resource needs and is typically used in OS‑level resource managers, while prevention is more common in application‑level concurrency control.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →