# Distributed Locks Best Practices: 10 Rules for Production Systems

> Master distributed locks best practices for production systems. Learn to implement them securely with Redis SET NX PX, unique IDs, Lua scripts, and short hold times to prevent deadlocks.

- Repository: [ByteByteGoHq/system-design-101](https://github.com/ByteByteGoHq/system-design-101)
- Tags: best-practices
- Published: 2026-02-28

---

**Implement distributed locks using Redis with atomic SET NX PX semantics, unique owner identifiers, Lua-scripted safe release, and minimal hold times to guarantee mutual exclusion without deadlocks.**

Distributed locks coordinate access to shared resources across multiple processes or services running on different machines. According to the ByteByteGoHq/system-design-101 repository, a well-designed lock must guarantee **mutual exclusion**, **fault tolerance**, and **low contention** without becoming a bottleneck. Below are the core best practices derived from the source code analysis.

## Choose the Right Lock Primitive

Use a primitive that provides atomic “set if not exists” semantics. In [`data/guides/how-can-redis-be-used.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/how-can-redis-be-used.md), Redis is highlighted as a primary mechanism for distributed locks, specifically recommending the `SET key value NX PX ttl` command pattern.

```javascript
const result = await redisClient.set(
  lockKey,
  lockId,
  { NX: true, PX: ttlMs }  // atomic set-if-not-exists + TTL
);

```

The `NX` flag ensures the key is set only if it does not exist, while `PX` sets the expiration time in milliseconds. This combination guarantees atomic acquisition and automatic expiration.

## Enforce Ownership with Unique Identifiers

Always include a unique client identifier, such as a UUID, as the lock value. This makes lock acquisition idempotent and ensures that only the owner can release the lock, preventing accidental unlocks by other processes.

```javascript
const lockId = crypto.randomUUID();  // unique owner id

```

Store this identifier alongside the lock key. When releasing, verify that the current holder matches the original acquirer before deletion.

## Implement Safe Release with Lua Scripts

Release the lock using a Lua script that checks the owner before deleting. As implemented in the ByteByteGo examples, this prevents a process from removing a lock that was reacquired by another node after the original TTL expired.

```lua
if redis.call("GET", KEYS[1]) == ARGV[1] then
  return redis.call("DEL", KEYS[1])
else
  return 0
end

```

Execute this atomically using `EVAL` to ensure the check-and-delete operation is not interrupted.

## Minimize Lock Hold Time and Contention

Keep the critical section as short as possible. The [`data/guides/pessimistic-vs-optimistic-locking.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/pessimistic-vs-optimistic-locking.md) guide emphasizes holding locks for the minimum possible time to reduce contention. Avoid I/O or long-running work while holding the lock.

Additionally, prefer fine-grained locks over global locks. The same guide recommends locking at the most granular level, such as rows rather than tables, to increase parallelism. Apply this principle by using per-entity or per-bucket lock keys rather than a single global lock.

```javascript
const lockKey = `inventory_lock:${productId}`;  // fine-grained per entity

```

## Build Resilient Acquisition and Release Patterns

Implement retry logic with exponential back-off and jitter when acquisition fails. This prevents "thundering herd" scenarios where many nodes repeatedly hammer the lock service simultaneously.

After a configurable number of attempts, fail gracefully rather than blocking indefinitely. Always release locks in a `finally` block to prevent deadlocks during exceptions.

```javascript
try {
  // critical section
} finally {
  await releaseLock(lockKey, lockId);  // guaranteed release
}

```

## Ensure High Availability and Observability

Deploy the lock service in a highly-available configuration, such as Redis Cluster or Sentinel, to survive node failures. A single-node lock store would defeat the purpose of distributed coordination.

Monitor lock acquisition latency, contention rate, and expiration events. The [`data/guides/why-do-we-need-to-use-a-distributed-lock.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/why-do-we-need-to-use-a-distributed-lock.md) file warns about duplicate execution and resource contention problems, which metrics can help detect early.

## Know When to Use Specialized Alternatives

Distributed locks are just one coordination tool. For leader election scenarios, consider consensus algorithms like Raft or ZooKeeper. For high-throughput rate limiting, use token-bucket algorithms instead of locks. As noted in [`data/guides/why-do-we-need-to-use-a-distributed-lock.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/why-do-we-need-to-use-a-distributed-lock.md), specific patterns often have specialized solutions that outperform generic locks.

## Production-Ready Implementation Example

Below is a complete lifecycle implementation combining the best practices from the ByteByteGoHq/system-design-101 repository.

```javascript
// Acquire lock --------------------------------------------------------------
async function acquireLock(key, ttlMs) {
  const lockId = crypto.randomUUID();               // unique owner id
  const result = await redisClient.set(
    key,
    lockId,
    { NX: true, PX: ttlMs }                         // atomic set-if-not-exists + TTL
  );
  return result === 'OK' ? lockId : null;           // null means lock not acquired
}

// Release lock --------------------------------------------------------------
async function releaseLock(key, lockId) {
  // Lua script guarantees we only delete if we still own the lock
  const script = `
    if redis.call("GET", KEYS[1]) == ARGV[1] then
      return redis.call("DEL", KEYS[1])
    else
      return 0
    end`;
  return await redisClient.eval(script, 1, key, lockId);
}

// Example critical section ---------------------------------------------------
async function updateInventory(productId, delta) {
  const lockKey = `inventory_lock:${productId}`;
  const lockId = await acquireLock(lockKey, 5000);   // 5-second lease
  if (!lockId) throw new Error('Could not acquire lock');

  try {
    // ----- critical section (keep short!) -----
    const stock = await redisClient.get(`stock:${productId}`);
    const newStock = Number(stock) + delta;
    await redisClient.set(`stock:${productId}`, newStock);
  } finally {
    await releaseLock(lockKey, lockId);             // always release
  }
}

```

## Summary

- **Use atomic SET NX PX**: Guarantee atomic acquisition and automatic expiration with Redis primitives.
- **Tag locks with UUIDs**: Ensure only the acquiring owner can release the lock.
- **Script the release**: Use Lua to check ownership before deletion, preventing race conditions.
- **Keep critical sections short**: Minimize hold time and use fine-grained keys to reduce contention.
- **Handle failures gracefully**: Implement retry with back-off and ensure locks release in `finally` blocks.
- **Deploy for HA**: Use Redis Cluster or Sentinel and monitor metrics to detect misbehavior.

## Frequently Asked Questions

### What is the primary risk of not using a TTL with distributed locks?

Without a time-to-live (TTL), a lock acquired by a process that subsequently crashes or loses network connectivity will never be released, causing a permanent deadlock. The [`data/guides/why-do-we-need-to-use-a-distributed-lock.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/why-do-we-need-to-use-a-distributed-lock.md) file highlights this as a critical failure scenario in distributed systems.

### Why must lock release be implemented as a Lua script in Redis?

A Lua script executes atomically on the Redis server, ensuring that the "check ownership then delete" operation cannot be interrupted by another process. Without this atomic guarantee, a different node could reacquire the lock after TTL expiration, then have its lock erroneously removed by the original owner's delayed release command.

### How does fine-grained locking improve system performance?

Fine-grained locking reduces contention by isolating critical sections to specific entities (e.g., individual product IDs) rather than using a single global lock. According to [`data/guides/pessimistic-vs-optimistic-locking.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/pessimistic-vs-optimistic-locking.md), locking at the most granular level allows parallel operations on different resources, significantly increasing throughput.

### When should I use ZooKeeper instead of Redis for distributed locks?

Use ZooKeeper or Raft-based solutions when you need strong consistency for leader election or when the coordination logic requires sequential node ordering and ephemeral nodes. While Redis provides high-performance locks, ZooKeeper offers stronger guarantees for complex consensus scenarios mentioned in [`data/guides/why-do-we-need-to-use-a-distributed-lock.md`](https://github.com/ByteByteGoHq/system-design-101/blob/main/data/guides/why-do-we-need-to-use-a-distributed-lock.md).