SPDLog Asynchronous Logging Performance: Architecture and Optimization Guide
SPDLog achieves high-performance asynchronous logging by decoupling the log call from I/O operations using a lock-free queue and dedicated thread pool, enabling 10–20× higher throughput than synchronous mode.
The gabime/spdlog repository implements a sophisticated asynchronous logging subsystem that minimizes latency on the application hot path. When using spdlog::async_logger, the library moves formatting and sink operations to background threads, reducing the critical path to a simple enqueue operation.
Architecture of SPDLog Asynchronous Logging
The asynchronous architecture centers on producer-consumer pattern separation. Application threads act as producers, while a dedicated thread pool handles consumption and output.
The Lock-Free Message Queue
At the core of the system sits the mpmc_blocking_q defined in include/spdlog/details/mpmc_blocking_q.h. This multi-producer-multi-consumer queue allows multiple application threads to enqueue log_msg objects without lock contention. The implementation uses atomic operations for enqueue/dequeue, ensuring that producer threads block only when the queue reaches capacity, not during the enqueue operation itself.
Thread Pool and Worker Architecture
The thread_pool class (declared in include/spdlog/details/thread_pool.h and implemented in include/spdlog/details/thread_pool-inl.h) manages a fixed pool of worker threads. Each worker follows this pattern:
- Dequeue a
log_msgfrom the lock-free queue - Format the message using the configured formatter (e.g.,
pattern_formatterfrominclude/spdlog/pattern_formatter.h) - Dispatch to the sink hierarchy (
include/spdlog/sinks/)
This separation ensures that expensive formatting and system calls (file writes, console output) never stall the application thread.
Performance Optimizations in the Source Code
SPDLog's asynchronous implementation incorporates several zero-copy and latency-reduction strategies.
Lock-Free MPMC Queue Design
The mpmc_blocking_q template in include/spdlog/details/mpmc_blocking_q.h eliminates mutex contention entirely. Unlike traditional blocking queues, this structure allows concurrent producers to push messages with only atomic compare-and-swap operations, scaling linearly with thread count under heavy load.
Thread-Local Buffer Management
The log_msg structure (defined in include/spdlog/details/log_msg.h) contains a small, pre-allocated buffer that avoids dynamic memory allocation during the logging call. For larger messages, the library streams content into a thread-local buffer before moving it into the queue, minimizing heap traffic and cache pollution on the producer thread.
Batch Processing and Flush Coalescing
Worker threads in the thread pool can batch multiple formatted messages before writing to disk or console. This strategy amortizes system call overhead across multiple log entries, particularly beneficial for file sinks where each write() syscall carries fixed overhead regardless of payload size.
Configurable Backpressure Policies
The async_overflow_policy enum (used in async_logger construction) allows tuning between latency and durability:
- block: Producer waits until queue space available (guarantees delivery)
- overrun_new: Newest message overwrites oldest when queue full (minimizes latency)
- overrun_old: Oldest messages dropped when queue full
Additionally, the init_thread_pool(queue_size, thread_count) API lets users match thread pool size to CPU core count and set queue depth based on burst tolerance requirements.
Implementing Asynchronous Loggers
To create an asynchronous logger, you must first initialize the global thread pool, then construct the logger with spdlog::async_logger.
Basic Async Logger Setup
The following pattern creates a thread pool with 4 workers and a queue size of 8192, then attaches a file sink:
#include <spdlog/async.h>
#include <spdlog/sinks/basic_file_sink.h>
int main() {
// Initialize thread pool: 8192 slots, 4 worker threads
spdlog::init_thread_pool(8192, 4);
// Create file sink
auto sink = std::make_shared<spdlog::sinks::basic_file_sink_mt>("async.log", true);
// Create async logger using the thread pool
auto async_logger = std::make_shared<spdlog::async_logger>(
"async_logger", sink, spdlog::thread_pool(),
spdlog::async_overflow_policy::block);
spdlog::register_logger(async_logger);
// Log 1 million messages without blocking application thread
for (int i = 0; i < 1'000'000; ++i) {
async_logger->info("Message {}", i);
}
spdlog::shutdown(); // Flushes pending messages before exit
}
Configuring Overflow Policies
To prevent blocking when the queue fills (useful for latency-sensitive applications), use the overrun_new policy:
auto async_logger = std::make_shared<spdlog::async_logger>(
"async_drop", sink, spdlog::thread_pool(),
spdlog::async_overflow_policy::overrun_new); // Newest message survives
Multi-Sink Asynchronous Logging
Multiple sinks can share the same thread pool, allowing simultaneous console and file output without double formatting:
auto console = std::make_shared<spdlog::sinks::stdout_color_sink_mt>();
auto file = std::make_shared<spdlog::sinks::basic_file_sink_mt>("combined.log");
std::vector<spdlog::sink_ptr> sinks{console, file};
auto combined_logger = std::make_shared<spdlog::async_logger>(
"combined", sinks.begin(), sinks.end(),
spdlog::thread_pool(),
spdlog::async_overflow_policy::block);
Performance Characteristics and Benchmarks
According to benchmarks in the repository's bench/ directory (specifically bench/utils.h), asynchronous mode demonstrates 10–20× higher throughput than synchronous logging under heavy concurrent load. The critical path for the application thread consists solely of:
- Constructing a
log_msgobject - Pushing it onto the lock-free queue
This O(1) operation typically costs only a few nanoseconds, keeping tail latency low even during I/O spikes. The dedicated thread pool absorbs formatting costs and disk write latency, preventing scheduler stalls in the application code.
Summary
- Lock-free queue: The
mpmc_blocking_qininclude/spdlog/details/mpmc_blocking_q.heliminates producer contention using atomic operations. - Thread pool:
thread_poolmanages background workers that handle formatting and I/O, decoupling these from the application hot path. - Zero-copy messages:
log_msguses pre-allocated buffers and thread-local storage to minimize heap allocations. - Tunable backpressure:
async_overflow_policylets you choose between blocking for durability or dropping messages for latency. - O(1) enqueue: Application threads pay only nanoseconds per log call, with actual output handled asynchronously.
Frequently Asked Questions
How does SPDLog's asynchronous mode achieve higher throughput than synchronous logging?
SPDLog's asynchronous mode achieves higher throughput by removing I/O bottlenecks from the application thread. In async_logger-inl.h, the log() method packages messages into log_msg objects and pushes them to the lock-free mpmc_blocking_q. A separate thread pool consumes these messages, performing formatting and sink writes on background threads. This allows the application to continue execution immediately after enqueueing, while the heavy lifting of disk writes and console output happens in parallel.
What is the difference between block and overrun_new overflow policies?
The block policy causes the calling thread to wait until queue space becomes available, guaranteeing message delivery at the cost of potential latency spikes. The overrun_new policy, defined in async_logger.h, instructs the queue to discard the oldest message when full, accepting the newest message instead. This prevents application stalls during traffic bursts but may result in lost log messages under sustained high load.
Can I use multiple sinks with a single asynchronous logger?
Yes, spdlog::async_logger accepts an iterator range of sinks in its constructor. When configured with multiple sinks (e.g., console and file), the thread pool worker formats the message once, then dispatches the result to all sinks in the hierarchy. This shared formatting reduces CPU overhead compared to creating separate loggers for each output destination.
How do I determine the optimal thread pool size and queue depth?
Size the thread pool to match your CPU core count or the number of concurrent I/O devices, typically set via init_thread_pool(queue_size, thread_count). Configure the queue size based on your burst tolerance—larger queues (e.g., 8192–65536 slots) absorb traffic spikes but consume more memory. If memory is constrained and occasional message loss is acceptable, use smaller queues with overrun_new policy to prevent blocking.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →