How Does mulle-objc Achieve Thread Safety Without a Global Lock: A Deep Dive into Fine-Grained Concurrency

Mulle-objc achieves thread safety without a global lock by combining atomic pointer operations for reference counting and method caches, per-classpair mutexes for initialization, thread-local storage for per-thread state, and isolated subsystem locks for ancillary data.

The mulle-objc runtime (mulle-objc/mulle-objc-runtime) was engineered from the ground up for highly concurrent environments. Instead of the traditional "big lock" approach that serializes all runtime operations, it employs a fine-grained, lock-free-where-possible strategy that minimizes contention and scales linearly across CPU cores.

Atomic Reference Counting With Zero Contention

Per-Object Atomic Operations

Every Objective-C object in mulle-objc stores its retain count directly in its header structure. Rather than protecting this count with a global mutex, the runtime uses atomic pointer operations that modify individual objects independently.

In src/mulle-objc-retain-release.h (lines 45-50), the _mulle_objc_object_increment_retaincount function performs lock-free increments:

static inline void _mulle_objc_object_increment_retaincount( void *obj )
{
    struct _mulle_objc_objectheader *header = _mulle_objc_object_get_objectheader( obj );
    // atomic increment – no global lock needed
    _mulle_atomic_pointer_increment( &header->_retaincount_1 );
}

This design ensures that retaining one object never blocks retaining another. The _mulle_atomic_pointer_increment and _mulle_atomic_pointer_decrement primitives, defined in src/mulle-objc-atomicpointer.h, provide hardware-level atomicity without requiring the runtime to serialize all memory operations through a single bottleneck.

Fine-Grained Locking for Class Initialization

Per-Classpair Mutex Strategy

Class initialization in mulle-objc uses localized locking that isolates contention to individual class pairs. When the runtime calls +initialize, it locks only the specific class being initialized, allowing other classes to initialize in parallel on different threads.

The implementation in src/mulle-objc-classpair.c (lines 55-58) initializes a dedicated mutex for each class pair:

mulle_thread_mutex_init( &pair->lock );

During initialization, as shown in src/mulle-objc-class-initialize.c, the runtime acquires this lock exclusively for the target class:

initialize_lock = _mulle_objc_classpair_get_lock( pair );
mulle_thread_mutex_lock( initialize_lock );   // only this class is locked

/* … set up super‑classes, caches, call +initialize … */

mulle_thread_mutex_unlock( initialize_lock );

This approach guarantees that initializing NSString does not block initializing NSArray, even when both operations occur simultaneously on separate threads.

Lock-Free Method Cache Updates

Atomic Cache Pointer Swapping

Method caches in mulle-objc support lock-free lookup by using atomic pointer indirection. When the runtime installs a new method cache, it performs an atomic write that instantly makes the new cache visible to all threads without requiring readers to acquire locks.

In src/mulle-objc-cache.h, the cache pointer is declared as mulle_atomic_pointer_t, enabling atomic operations like:

/* install a freshly built cache atomically */
_mulle_atomic_pointer_write( &cls->cache, new_cache );

Readers use _mulle_atomic_pointer_read to access cache entries, ensuring they see either the complete old cache or the complete new cache, never a partially constructed state. This technique eliminates cache lookup latency caused by mutex contention during cache updates.

Impcache Atomic Indirection

For fast method tables (impcaches), the runtime employs similar atomic strategies. The _mulle_objc_impcachepivot_get_entries_atomic function in src/mulle-objc-impcache.h allows the entire method table to be replaced in a single atomic step, keeping method dispatch lock-free in the common case.

Thread-Local Storage for Per-Thread State

Eliminating Cross-Thread Contention

Mulle-objc avoids sharing mutable thread state by storing per-thread data in thread-local storage (TLS). Each OS thread maintains its own struct _mulle_objc_threadinfo, eliminating the need for locks when accessing thread-specific counters or flags.

In src/mulle-objc-universe.c (lines 737-739), the universe creates a TLS key during initialization:

if (mulle_thread_tss_create( (void *)mulle_objc_threadinfo_free,
                             &universe->threadkey ))
    mulle_objc_universe_fail_perror( NULL, "No more thread local storage keys" );

When a thread accesses the runtime, it retrieves its private threadinfo structure via mulle_thread_tss_get. Because no other thread can access this memory, operations like incrementing the thread_counter require no synchronization primitives, removing a significant source of cache-line bouncing in multi-threaded applications.

Isolated Subsystem Locks

Dedicated Mutexes for Ancillary Data

While the hot paths remain lock-free, mulle-objc maintains tiny, dedicated mutexes for specific subsystems that require synchronization. These locks never appear in object allocation, retention, or message dispatch paths.

In src/mulle-objc-universe.c (lines 748-752), the universe initializes three isolated locks:

mulle_thread_mutex_init( &universe->debug.lock );
mulle_thread_mutex_init( &universe->waitqueues.lock );
mulle_thread_mutex_init( &universe->lock );

The debug.lock protects trace and debug state, waitqueues.lock manages thread waiting queues, and universe->lock guards configuration flags. By segregating these concerns, the runtime ensures that debugging operations or configuration changes cannot stall the message-passing pipeline.

Summary

  • Atomic reference counting uses per-object _mulle_atomic_pointer_increment operations in mulle-objc-retain-release.h, eliminating global lock contention during retain/release cycles.
  • Per-classpair mutexes in mulle-objc-classpair.c isolate class initialization, allowing parallel +initialize execution across different classes.
  • Lock-free method caches leverage atomic pointer writes in mulle-objc-cache.h, enabling thread-safe cache updates without blocking message dispatch.
  • Thread-local storage via mulle_thread_tss_create in mulle-objc-universe.c gives each thread private state, removing cross-thread contention for thread-specific data.
  • Subsystem-specific locks protect only debugging and wait-queue structures, keeping the critical object messaging path completely lock-free.

Frequently Asked Questions

What synchronization primitive does mulle-objc use for reference counting?

Mulle-objc uses atomic pointer operations via _mulle_atomic_pointer_increment and _mulle_atomic_pointer_decrement defined in src/mulle-objc-atomicpointer.h. These hardware-level atomic instructions modify an object's retain count in its header without requiring locks, ensuring that reference counting on one object never interferes with operations on another object.

How does mulle-objc prevent deadlocks during class initialization?

The runtime assigns a dedicated mutex per classpair (pair->lock in src/mulle-objc-classpair.c). When initializing a class, the runtime locks only that specific class's mutex, avoiding the circular dependencies that often plague global-lock designs. Since each class pair has its own lock, initializing class A cannot block waiting for class B's initialization lock unless there is an actual superclass relationship requiring ordered initialization.

Can method cache updates occur concurrently with message dispatch?

Yes. Method cache updates use atomic pointer writes (_mulle_atomic_pointer_write in src/mulle-objc-cache.h) that swap cache pointers atomically. Message dispatch reads the cache pointer atomically with _mulle_atomic_pointer_read, ensuring dispatch either sees the old valid cache or the new valid cache, never a partially constructed cache. This lock-free approach allows cache growth and optimization to occur without pausing message sending.

Why doesn't mulle-objc use a global lock for the entire runtime?

A global lock would serialize all runtime operations, creating a bottleneck on multi-core systems. Instead, mulle-objc follows a fine-grained concurrency model where mutable state is either protected by object-local atomics, class-local mutexes, or stored in thread-local storage. According to the source code in src/mulle-objc-universe.c, only ancillary subsystems like debugging and wait-queues use small dedicated locks, leaving the critical paths of object allocation, retention, and message dispatch lock-free and highly scalable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →