How Method Inlining Works in mulle-objc-runtime for Performance-Critical Code

Method inlining in mulle-objc-runtime eliminates function call overhead by embedding the inline cache lookup directly into the caller's compiled code, using static-inline helpers defined in src/mulle-objc-call.h to accelerate hot paths like retain-release loops and tight message dispatches.

The mulle-objc-runtime provides a high-performance Objective-C implementation optimized for modern CPU architectures. Understanding how method inlining works for performance-critical code reveals why this runtime outperforms traditional dispatch mechanisms in tight loops and high-frequency reference-counting scenarios.

Why Method Inlining Matters for Performance

The runtime uses an inline cache (ICache) to accelerate Objective-C method dispatch. When a message is sent, the runtime first attempts to locate the implementation (IMP) in a per-class cache. If the cache entry is present, the call resolves without a full lookup, saving both branch mispredictions and memory indirections.

For hot paths—such as tight numeric loops, retain/release churn, or frequent message sends—the overhead of a regular function call remains noticeable. To eliminate this cost, the runtime provides static-inline helpers that embed the cache walk directly into the caller's compiled code.

Key benefits of this approach include:

  • Zero-overhead calls – The compiler emits the exact sequence of loads, masks, and branch checks directly at the call site.
  • Eliminated stack frames – No extra function prologue or epilogue; the caller's frame is reused.
  • Improved branch prediction – Hot-path code remains in a tight loop, allowing the CPU to maintain prediction state.
  • Fast-method-table compatibility – When __MULLE_OBJC_FCS__ is defined, inline helpers can fall back to the fast-method-table without an extra function call.

Core Inline Helpers in src/mulle-objc-call.h

The primary inline dispatch utilities live in src/mulle-objc-call.h and are marked with MULLE_C_STATIC_ALWAYS_INLINE to force compiler inlining:

Inline helper Purpose Source location
mulle_objc_object_call_inline Public entry point when the selector is known at compile time and the object is non-null; calls the full inline dispatcher. src/mulle-objc-call.h#L25-L35
_mulle_objc_object_get_imp_inline_full_no_fcs Returns the cached IMP without fast-method-table (FCS) handling; used by the regular non-inline call. src/mulle-objc-call.h#L38-L52
_mulle_objc_object_call_inline_full Full inlined dispatch that may invoke the fast-method-table when __MULLE_OBJC_FCS__ is defined; returns the method call result. src/mulle-objc-call.h#L96-L112
_mulle_objc_object_retain_inline / _mulle_objc_object_release_inline Thin wrappers around reference-counting helpers for performance-critical loops. src/mulle-objc-retain-release.c#L213-L224

The Inline Cache Walk Implementation

The actual cache walk implementation performs four critical steps directly in the generated machine code:

  1. Cache pivot fetch – Retrieves cache entries via _mulle_objc_cachepivot_get_entries_atomic(&cls->cachepivot.pivot).
  2. Mask extraction – Extracts the probing mask from cache->mask.
  3. Probing loop – Repeatedly calculates offset = offset & mask and loads entries until finding a matching key or an empty slot.
  4. Miss handling – On cache miss, returns the miss callback from the ICache (icache->callback.call_cache_miss).

This entire sequence emits inline into the caller, resulting in a tight sequence of loads, logical AND operations, and at most one conditional branch per iteration.

Interaction with the ICache

When the inline walk encounters a miss (if (! entry->key.uniqueid)), it fetches the miss handler from the cache's icache structure without breaking the inline paradigm:

icache = _mulle_objc_cache_get_impcache_from_cache(cache);
f = icache->callback.call_cache_miss;

The call_cache_miss function, defined in src/mulle-objc-impcache.c, performs the heavyweight lookup via _mulle_objc_class_lookup_superimplementation_inline_nofail before patching the cache entry. Because this heavy path executes only on misses, the fast path (cache hit) remains completely inlined, preserving maximum throughput for the common case.

When to Use the Inline API

Choose the appropriate dispatch method based on your execution context:

  • Hot tight loops (e.g., array processing) with compile-time known selectors – Use mulle_objc_object_call_inline(obj, SEL, param).
  • Retain/release hot paths – Use _mulle_objc_object_retain_inline(obj) or _mulle_objc_object_release_inline(obj) to minimize overhead.
  • Variable selectors – Avoid inlining; use the non-inline dispatcher mulle_objc_object_call, which safely handles dynamic selectors.
  • Debug builds – Prefer non-inline helpers to maintain readable stack traces; inline versions deliberately skip frame generation.

Practical Code Examples

Retain-Release in Tight Loops

The test suite in test/demo/retain-release.c demonstrates eliminating call overhead in high-frequency reference counting:

// Performance-critical loop
for (int i = 0; i < 1000000; ++i)
{
    _mulle_objc_object_retain_inline(obj);
    _mulle_objc_object_release_inline(obj);
}

This compiles to a single load + atomic increment/decrement sequence per iteration, compared to the extra function call required by the public API.

Inlined Message Dispatch

For hot message sends with known selectors:

extern mulle_objc_methodid_t MULLE_OBJC_FOO_METHODID;

void *result = mulle_objc_object_call_inline(my_obj, 
                                             MULLE_OBJC_FOO_METHODID, 
                                             NULL);

When the cache contains the implementation, the generated assembly executes a direct sequence of memory loads and a single branch without intermediate function calls:

mov    rax, [my_obj]               ; load isa
mov    rdx, [rax + cachepivot]      ; cache pivot
mov    rcx, [rdx]                   ; entries pointer
and    rsi, mask                    ; methodid & mask
mov    r8,  [rcx + rsi]            ; entry
cmp    [r8].uniqueid, methodid
jne    miss_handler
mov    rax, [r8].imp                ; cached IMP
call   rax

How Inline Helpers Stay Synchronized

Because helpers are defined as MULLE_C_STATIC_ALWAYS_INLINE, the compiler must inline them at the point of use during optimization. Any changes to the cache layout in src/mulle-objc-call.h automatically propagate to all call sites at compile time, ensuring the inline code always matches the current runtime structures without manual synchronization.

Summary

  • Method inlining embeds the cache lookup logic directly into caller code via MULLE_C_STATIC_ALWAYS_INLINE functions in src/mulle-objc-call.h.
  • The inline cache walk performs pivot fetching, mask extraction, and linear probing without function call overhead.
  • Cache misses delegate to call_cache_miss in src/mulle-objc-impcache.c, keeping the fast path entirely inline.
  • Use mulle_objc_object_call_inline for compile-time known selectors and _mulle_objc_object_retain_inline for high-frequency reference counting.
  • Inline helpers automatically reflect runtime structure changes at compilation, eliminating version skew risks.

Frequently Asked Questions

What is the difference between mulle_objc_object_call_inline and the regular dispatch function?

mulle_objc_object_call_inline requires the selector to be known at compile time and embeds the entire cache lookup directly into your function's machine code, eliminating the call overhead to the dispatcher. The regular mulle_objc_object_call performs the lookup through a non-inline function call in src/mulle-objc-call.c, which adds stack frame overhead but handles dynamic selectors safely.

When should I avoid using inline method dispatch?

Avoid inline dispatch when the selector is determined at runtime (preventing compiler optimization), when building debug configurations where stack trace readability matters, or when calling from dynamically loaded plugins that may not share the same inline cache assumptions. In these cases, the non-inline API in src/mulle-objc-call.c provides safer indirection.

How does the inline cache handle cache misses without losing performance?

On cache miss, the inline code fetches icache->callback.call_cache_miss from the inline cache structure and jumps to the miss handler defined in src/mulle-objc-impcache.c. This handler performs the expensive method lookup and cache update, but because misses are statistically rare in hot loops, the common-case cache hit remains inlined and executes at maximum speed without function call overhead.

Does method inlining work with the fast-method-table (FCS) optimization?

Yes. When __MULLE_OBJC_FCS__ is defined, _mulle_objc_object_call_inline_full automatically includes the fast-method-table fallback path inline. This allows the code to check the fast-method-table without an additional function call, maintaining the zero-overhead characteristic even when utilizing this additional optimization layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →