# How Does mulle-objc Achieve Thread Safety Without a Global Lock: A Deep Dive into Fine-Grained Concurrency

> Discover how mulle-objc uses atomic operations, per-classpair mutexes, thread-local storage, and subsystem locks for fine-grained concurrency and thread safety without a global lock.

- Repository: [mulle-objc/mulle-objc-runtime](https://github.com/mulle-objc/mulle-objc-runtime)
- Tags: deep-dive
- Published: 2026-03-07

---

**Mulle-objc achieves thread safety without a global lock by combining atomic pointer operations for reference counting and method caches, per-classpair mutexes for initialization, thread-local storage for per-thread state, and isolated subsystem locks for ancillary data.**

The mulle-objc runtime (`mulle-objc/mulle-objc-runtime`) was engineered from the ground up for highly concurrent environments. Instead of the traditional "big lock" approach that serializes all runtime operations, it employs a **fine-grained, lock-free-where-possible** strategy that minimizes contention and scales linearly across CPU cores.

## Atomic Reference Counting With Zero Contention

### Per-Object Atomic Operations

Every Objective-C object in mulle-objc stores its retain count directly in its header structure. Rather than protecting this count with a global mutex, the runtime uses **atomic pointer operations** that modify individual objects independently.

In [`src/mulle-objc-retain-release.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-retain-release.h) (lines 45-50), the `_mulle_objc_object_increment_retaincount` function performs lock-free increments:

```c
static inline void _mulle_objc_object_increment_retaincount( void *obj )
{
    struct _mulle_objc_objectheader *header = _mulle_objc_object_get_objectheader( obj );
    // atomic increment – no global lock needed
    _mulle_atomic_pointer_increment( &header->_retaincount_1 );
}

```

This design ensures that retaining one object never blocks retaining another. The `_mulle_atomic_pointer_increment` and `_mulle_atomic_pointer_decrement` primitives, defined in [`src/mulle-objc-atomicpointer.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-atomicpointer.h), provide hardware-level atomicity without requiring the runtime to serialize all memory operations through a single bottleneck.

## Fine-Grained Locking for Class Initialization

### Per-Classpair Mutex Strategy

Class initialization in mulle-objc uses **localized locking** that isolates contention to individual class pairs. When the runtime calls `+initialize`, it locks only the specific class being initialized, allowing other classes to initialize in parallel on different threads.

The implementation in [`src/mulle-objc-classpair.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-classpair.c) (lines 55-58) initializes a dedicated mutex for each class pair:

```c
mulle_thread_mutex_init( &pair->lock );

```

During initialization, as shown in [`src/mulle-objc-class-initialize.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-class-initialize.c), the runtime acquires this lock exclusively for the target class:

```c
initialize_lock = _mulle_objc_classpair_get_lock( pair );
mulle_thread_mutex_lock( initialize_lock );   // only this class is locked

/* … set up super‑classes, caches, call +initialize … */

mulle_thread_mutex_unlock( initialize_lock );

```

This approach guarantees that initializing `NSString` does not block initializing `NSArray`, even when both operations occur simultaneously on separate threads.

## Lock-Free Method Cache Updates

### Atomic Cache Pointer Swapping

Method caches in mulle-objc support **lock-free lookup** by using atomic pointer indirection. When the runtime installs a new method cache, it performs an atomic write that instantly makes the new cache visible to all threads without requiring readers to acquire locks.

In [`src/mulle-objc-cache.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-cache.h), the cache pointer is declared as `mulle_atomic_pointer_t`, enabling atomic operations like:

```c
/* install a freshly built cache atomically */
_mulle_atomic_pointer_write( &cls->cache, new_cache );

```

Readers use `_mulle_atomic_pointer_read` to access cache entries, ensuring they see either the complete old cache or the complete new cache, never a partially constructed state. This technique eliminates cache lookup latency caused by mutex contention during cache updates.

### Impcache Atomic Indirection

For fast method tables (impcaches), the runtime employs similar atomic strategies. The `_mulle_objc_impcachepivot_get_entries_atomic` function in [`src/mulle-objc-impcache.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-impcache.h) allows the entire method table to be replaced in a single atomic step, keeping method dispatch lock-free in the common case.

## Thread-Local Storage for Per-Thread State

### Eliminating Cross-Thread Contention

Mulle-objc avoids sharing mutable thread state by storing per-thread data in **thread-local storage (TLS)**. Each OS thread maintains its own `struct _mulle_objc_threadinfo`, eliminating the need for locks when accessing thread-specific counters or flags.

In [`src/mulle-objc-universe.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-universe.c) (lines 737-739), the universe creates a TLS key during initialization:

```c
if (mulle_thread_tss_create( (void *)mulle_objc_threadinfo_free,
                             &universe->threadkey ))
    mulle_objc_universe_fail_perror( NULL, "No more thread local storage keys" );

```

When a thread accesses the runtime, it retrieves its private `threadinfo` structure via `mulle_thread_tss_get`. Because no other thread can access this memory, operations like incrementing the `thread_counter` require no synchronization primitives, removing a significant source of cache-line bouncing in multi-threaded applications.

## Isolated Subsystem Locks

### Dedicated Mutexes for Ancillary Data

While the hot paths remain lock-free, mulle-objc maintains **tiny, dedicated mutexes** for specific subsystems that require synchronization. These locks never appear in object allocation, retention, or message dispatch paths.

In [`src/mulle-objc-universe.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-universe.c) (lines 748-752), the universe initializes three isolated locks:

```c
mulle_thread_mutex_init( &universe->debug.lock );
mulle_thread_mutex_init( &universe->waitqueues.lock );
mulle_thread_mutex_init( &universe->lock );

```

The `debug.lock` protects trace and debug state, `waitqueues.lock` manages thread waiting queues, and `universe->lock` guards configuration flags. By segregating these concerns, the runtime ensures that debugging operations or configuration changes cannot stall the message-passing pipeline.

## Summary

- **Atomic reference counting** uses per-object `_mulle_atomic_pointer_increment` operations in [`mulle-objc-retain-release.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/mulle-objc-retain-release.h), eliminating global lock contention during retain/release cycles.
- **Per-classpair mutexes** in [`mulle-objc-classpair.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/mulle-objc-classpair.c) isolate class initialization, allowing parallel `+initialize` execution across different classes.
- **Lock-free method caches** leverage atomic pointer writes in [`mulle-objc-cache.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/mulle-objc-cache.h), enabling thread-safe cache updates without blocking message dispatch.
- **Thread-local storage** via `mulle_thread_tss_create` in [`mulle-objc-universe.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/mulle-objc-universe.c) gives each thread private state, removing cross-thread contention for thread-specific data.
- **Subsystem-specific locks** protect only debugging and wait-queue structures, keeping the critical object messaging path completely lock-free.

## Frequently Asked Questions

### What synchronization primitive does mulle-objc use for reference counting?

Mulle-objc uses **atomic pointer operations** via `_mulle_atomic_pointer_increment` and `_mulle_atomic_pointer_decrement` defined in [`src/mulle-objc-atomicpointer.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-atomicpointer.h). These hardware-level atomic instructions modify an object's retain count in its header without requiring locks, ensuring that reference counting on one object never interferes with operations on another object.

### How does mulle-objc prevent deadlocks during class initialization?

The runtime assigns a **dedicated mutex per classpair** (`pair->lock` in [`src/mulle-objc-classpair.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-classpair.c)). When initializing a class, the runtime locks only that specific class's mutex, avoiding the circular dependencies that often plague global-lock designs. Since each class pair has its own lock, initializing class A cannot block waiting for class B's initialization lock unless there is an actual superclass relationship requiring ordered initialization.

### Can method cache updates occur concurrently with message dispatch?

Yes. Method cache updates use **atomic pointer writes** (`_mulle_atomic_pointer_write` in [`src/mulle-objc-cache.h`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-cache.h)) that swap cache pointers atomically. Message dispatch reads the cache pointer atomically with `_mulle_atomic_pointer_read`, ensuring dispatch either sees the old valid cache or the new valid cache, never a partially constructed cache. This lock-free approach allows cache growth and optimization to occur without pausing message sending.

### Why doesn't mulle-objc use a global lock for the entire runtime?

A global lock would serialize all runtime operations, creating a bottleneck on multi-core systems. Instead, mulle-objc follows a **fine-grained concurrency** model where mutable state is either protected by object-local atomics, class-local mutexes, or stored in thread-local storage. According to the source code in [`src/mulle-objc-universe.c`](https://github.com/mulle-objc/mulle-objc-runtime/blob/main/src/mulle-objc-universe.c), only ancillary subsystems like debugging and wait-queues use small dedicated locks, leaving the critical paths of object allocation, retention, and message dispatch lock-free and highly scalable.