SPDLog Asynchronous Logging Performance: Architecture and Optimization Guide

SPDLog achieves high-performance asynchronous logging by decoupling the log call from I/O operations using a lock-free queue and dedicated thread pool, enabling 10–20× higher throughput than synchronous mode.

The gabime/spdlog repository implements a sophisticated asynchronous logging subsystem that minimizes latency on the application hot path. When using spdlog::async_logger, the library moves formatting and sink operations to background threads, reducing the critical path to a simple enqueue operation.

Architecture of SPDLog Asynchronous Logging

The asynchronous architecture centers on producer-consumer pattern separation. Application threads act as producers, while a dedicated thread pool handles consumption and output.

The Lock-Free Message Queue

At the core of the system sits the mpmc_blocking_q defined in include/spdlog/details/mpmc_blocking_q.h. This multi-producer-multi-consumer queue allows multiple application threads to enqueue log_msg objects without lock contention. The implementation uses atomic operations for enqueue/dequeue, ensuring that producer threads block only when the queue reaches capacity, not during the enqueue operation itself.

Thread Pool and Worker Architecture

The thread_pool class (declared in include/spdlog/details/thread_pool.h and implemented in include/spdlog/details/thread_pool-inl.h) manages a fixed pool of worker threads. Each worker follows this pattern:

  1. Dequeue a log_msg from the lock-free queue
  2. Format the message using the configured formatter (e.g., pattern_formatter from include/spdlog/pattern_formatter.h)
  3. Dispatch to the sink hierarchy (include/spdlog/sinks/)

This separation ensures that expensive formatting and system calls (file writes, console output) never stall the application thread.

Performance Optimizations in the Source Code

SPDLog's asynchronous implementation incorporates several zero-copy and latency-reduction strategies.

Lock-Free MPMC Queue Design

The mpmc_blocking_q template in include/spdlog/details/mpmc_blocking_q.h eliminates mutex contention entirely. Unlike traditional blocking queues, this structure allows concurrent producers to push messages with only atomic compare-and-swap operations, scaling linearly with thread count under heavy load.

Thread-Local Buffer Management

The log_msg structure (defined in include/spdlog/details/log_msg.h) contains a small, pre-allocated buffer that avoids dynamic memory allocation during the logging call. For larger messages, the library streams content into a thread-local buffer before moving it into the queue, minimizing heap traffic and cache pollution on the producer thread.

Batch Processing and Flush Coalescing

Worker threads in the thread pool can batch multiple formatted messages before writing to disk or console. This strategy amortizes system call overhead across multiple log entries, particularly beneficial for file sinks where each write() syscall carries fixed overhead regardless of payload size.

Configurable Backpressure Policies

The async_overflow_policy enum (used in async_logger construction) allows tuning between latency and durability:

  • block: Producer waits until queue space available (guarantees delivery)
  • overrun_new: Newest message overwrites oldest when queue full (minimizes latency)
  • overrun_old: Oldest messages dropped when queue full

Additionally, the init_thread_pool(queue_size, thread_count) API lets users match thread pool size to CPU core count and set queue depth based on burst tolerance requirements.

Implementing Asynchronous Loggers

To create an asynchronous logger, you must first initialize the global thread pool, then construct the logger with spdlog::async_logger.

Basic Async Logger Setup

The following pattern creates a thread pool with 4 workers and a queue size of 8192, then attaches a file sink:

#include <spdlog/async.h>
#include <spdlog/sinks/basic_file_sink.h>

int main() {
    // Initialize thread pool: 8192 slots, 4 worker threads
    spdlog::init_thread_pool(8192, 4);

    // Create file sink
    auto sink = std::make_shared<spdlog::sinks::basic_file_sink_mt>("async.log", true);

    // Create async logger using the thread pool
    auto async_logger = std::make_shared<spdlog::async_logger>(
        "async_logger", sink, spdlog::thread_pool(),
        spdlog::async_overflow_policy::block);

    spdlog::register_logger(async_logger);

    // Log 1 million messages without blocking application thread
    for (int i = 0; i < 1'000'000; ++i) {
        async_logger->info("Message {}", i);
    }

    spdlog::shutdown();   // Flushes pending messages before exit
}

Configuring Overflow Policies

To prevent blocking when the queue fills (useful for latency-sensitive applications), use the overrun_new policy:

auto async_logger = std::make_shared<spdlog::async_logger>(
    "async_drop", sink, spdlog::thread_pool(),
    spdlog::async_overflow_policy::overrun_new); // Newest message survives

Multi-Sink Asynchronous Logging

Multiple sinks can share the same thread pool, allowing simultaneous console and file output without double formatting:

auto console = std::make_shared<spdlog::sinks::stdout_color_sink_mt>();
auto file = std::make_shared<spdlog::sinks::basic_file_sink_mt>("combined.log");
std::vector<spdlog::sink_ptr> sinks{console, file};

auto combined_logger = std::make_shared<spdlog::async_logger>(
    "combined", sinks.begin(), sinks.end(),
    spdlog::thread_pool(),
    spdlog::async_overflow_policy::block);

Performance Characteristics and Benchmarks

According to benchmarks in the repository's bench/ directory (specifically bench/utils.h), asynchronous mode demonstrates 10–20× higher throughput than synchronous logging under heavy concurrent load. The critical path for the application thread consists solely of:

  1. Constructing a log_msg object
  2. Pushing it onto the lock-free queue

This O(1) operation typically costs only a few nanoseconds, keeping tail latency low even during I/O spikes. The dedicated thread pool absorbs formatting costs and disk write latency, preventing scheduler stalls in the application code.

Summary

  • Lock-free queue: The mpmc_blocking_q in include/spdlog/details/mpmc_blocking_q.h eliminates producer contention using atomic operations.
  • Thread pool: thread_pool manages background workers that handle formatting and I/O, decoupling these from the application hot path.
  • Zero-copy messages: log_msg uses pre-allocated buffers and thread-local storage to minimize heap allocations.
  • Tunable backpressure: async_overflow_policy lets you choose between blocking for durability or dropping messages for latency.
  • O(1) enqueue: Application threads pay only nanoseconds per log call, with actual output handled asynchronously.

Frequently Asked Questions

How does SPDLog's asynchronous mode achieve higher throughput than synchronous logging?

SPDLog's asynchronous mode achieves higher throughput by removing I/O bottlenecks from the application thread. In async_logger-inl.h, the log() method packages messages into log_msg objects and pushes them to the lock-free mpmc_blocking_q. A separate thread pool consumes these messages, performing formatting and sink writes on background threads. This allows the application to continue execution immediately after enqueueing, while the heavy lifting of disk writes and console output happens in parallel.

What is the difference between block and overrun_new overflow policies?

The block policy causes the calling thread to wait until queue space becomes available, guaranteeing message delivery at the cost of potential latency spikes. The overrun_new policy, defined in async_logger.h, instructs the queue to discard the oldest message when full, accepting the newest message instead. This prevents application stalls during traffic bursts but may result in lost log messages under sustained high load.

Can I use multiple sinks with a single asynchronous logger?

Yes, spdlog::async_logger accepts an iterator range of sinks in its constructor. When configured with multiple sinks (e.g., console and file), the thread pool worker formats the message once, then dispatches the result to all sinks in the hierarchy. This shared formatting reduces CPU overhead compared to creating separate loggers for each output destination.

How do I determine the optimal thread pool size and queue depth?

Size the thread pool to match your CPU core count or the number of concurrent I/O devices, typically set via init_thread_pool(queue_size, thread_count). Configure the queue size based on your burst tolerance—larger queues (e.g., 8192–65536 slots) absorb traffic spikes but consume more memory. If memory is constrained and occasional message loss is acceptable, use smaller queues with overrun_new policy to prevent blocking.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →