How fmtlib Parses Format Specifiers: A Deep Dive into {fmt}'s State-Machine Parser
{fmt} parses format specifiers using a constexpr state machine in include/fmt/core.h, driven by the parse_format_specs function that processes alignment, sign, width, precision, and type components in strict order.
The {fmt} library implements one of the most robust format specifier parsers in modern C++. Unlike legacy printf parsing, {fmt} performs compile-time validation where possible, using a carefully structured state machine to ensure specifiers are syntactically correct and semantically appropriate for the argument type. This article examines the actual implementation in the fmtlib/fmt repository to explain how format specifier parsing works under the hood.
The Entry Point: parse_format_specs
All format specifier parsing begins in include/fmt/core.h at line 55, where the parse_format_specs function is defined:
auto parse_format_specs(const Char* begin, const Char* end,
dynamic_format_specs<Char>& specs, ...)
This constexpr function iterates through characters between { and }, building a format_specs object that stores every parsed component. The function returns an iterator positioned at the closing } or the end of the specifier string.
If the format string contains an empty specifier {}, the parser returns immediately without processing. Otherwise, it initializes a state machine to track which components have been encountered.
State Machine Architecture
The parser uses an internal state enum to enforce correct ordering of specifier components. The states progress through: start → align → sign → hash → zero → width → precision → locale.
The helper function enter_state (lines 70-77) validates transitions:
// From include/fmt/core.h
auto enter_state = [&](state new_state) {
if (int(state) >= int(new_state)) report_error("invalid format spec");
state = new_state;
};
This strict ordering prevents malformed specifiers like {:10<} (width before alignment) from compiling or running.
Component-by-Component Parsing
Alignment and Fill Characters
Alignment tokens <, >, and ^ are detected at lines 94-101. The parser supports multibyte fill characters—if a character precedes an alignment token, it is extracted as the fill sequence:
// Lines 71-88: Extracting fill character
if (*begin != '{' && *begin != '}') {
auto c = *begin++; // consume potential fill character
if (begin != end && *begin == '<' || *begin == '>' || *begin == '^') {
specs.set_fill(c); // line 84
// ... set alignment based on *begin
}
}
This enables patterns like {:*>10} for right-aligned, asterisk-filled output.
Sign, Alternative Form, and Zero-Padding
- Sign (
+,-, space): Lines 102-109, restricted to signed integers and floating-point types viain(arg_type, sint_set | float_set) - Alternative form (
#): Lines 111-114, enabling0xprefixes for hex, etc. - Zero-padding (
0): Lines 115-124, treated as numeric alignment with fill'0', rejected for non-numeric arguments
Each modifier validates the argument type before acceptance, preventing nonsensical combinations like {:+} on a string.
Width and Precision
Width is parsed by parse_width (lines 130-132), handling both literal values and dynamic specifications ({} or {n}):
specs.set_dynamic_width(arg_id, index); // line 133
Precision follows a dot (.) at lines 136-144, using parse_precision with similar dynamic value support (specs.set_dynamic_precision at line 149).
Locale Flag and Presentation Type
The L flag (lines 140-142) enables locale-aware formatting for numeric separators. Finally, the presentation type (lines 146-166) determines output format through parse_presentation_type, which validates against type-specific allowed sets:
| Type Character | Meaning | Allowed Argument Types |
|---|---|---|
d, i |
Decimal integer | Integral types |
x, X |
Hexadecimal | Integral types |
o |
Octal | Integral types |
b, B |
Binary | Integral types |
f, F |
Fixed floating-point | Floating-point types |
e, E |
Scientific notation | Floating-point types |
g, G |
General format | Floating-point types |
a, A |
Hexadecimal floating-point | Floating-point types |
c |
Character | Integral/character types |
s |
String | String types |
p |
Pointer | Pointer types |
? |
Debug format | Any type with debug support |
The format_specs Storage Structure
Parsed data is stored in struct format_specs : basic_specs (line 831), which inherits core flags and adds:
- Fill character (multibyte support)
- Dynamic width/precision tracking
- Argument ID references for runtime values
This structure is consumed by formatter write overloads in include/fmt/format.h to produce final output.
Practical Examples
These examples exercise the complete specifier parser:
#include <fmt/core.h>
#include <fmt/format.h>
int main() {
int v = 42;
double pi = 3.14159;
// Fill, alignment, width
fmt::print("{:*^12}\n", v); // *****42*****
// Sign, zero-fill, width
fmt::print("{:+08}\n", v); // +0000042
// Alternative form, hex type
fmt::print("{:#x}\n", v); // 0x2a
// Precision and floating-point type
fmt::print("{:.2f}\n", pi); // 3.14
// Locale-aware formatting
fmt::print("{:L}\n", 1234); // 1,234 (locale-dependent)
// Debug type for pairs
fmt::print("{:?}\n", std::make_pair(1, 2)); // (1, 2)
}
Key Source Files
Understanding format specifier parsing requires familiarity with these files in the fmtlib/fmt repository:
include/fmt/core.h— Core definitions includingformat_specs,basic_specs, and the completeparse_format_specsimplementationinclude/fmt/format.h— Entry point for formatting; forwards parsed specs to writer functionsinclude/fmt/format-inl.h— Inline utilities that consumeformat_specsfor actual output generation
Summary
parse_format_specsininclude/fmt/core.his the central parsing function, operating at compile-time where possible- A strict state machine enforces correct specifier component ordering through the
stateenum andenter_statehelper - Type validation occurs at every stage, preventing invalid combinations like sign flags on strings
- Dynamic specifiers (
{}for width/precision) are tracked separately for runtime resolution - Results are stored in
format_specsobjects consumed by writer functions in the format implementation
Frequently Asked Questions
Why does {fmt} use a state machine for parsing format specifiers?
The state machine design enforces that format specifier components appear in a strict, unambiguous order. This prevents malformed patterns like {:10<} where width precedes alignment, and enables clear error messaging. The enter_state helper at lines 70-77 validates transitions by comparing enum values, catching errors at compile time for literal format strings.
Can I use multibyte Unicode characters as fill characters in {fmt}?
Yes. The parser at lines 71-88 handles fill character extraction before checking for alignment tokens. The specs.set_fill call stores the complete character sequence, supporting Unicode fill characters when using wchar_t or char8_t/char16_t/char32_t specializations.
How does {fmt} validate that a type character matches the argument type?
Each presentation type case calls parse_presentation_type with an allowed type set (lines 146-166). For example, case 'x': return parse_presentation_type(pres::hex, integral_set) restricts hex formatting to integral types. The function checks in(arg_type, allowed_set) and reports errors for mismatches, providing type-safe formatting that printf cannot match.
What happens when width or precision is specified dynamically with {}?
The parser detects the opening brace in parse_width or parse_precision, records the dynamic nature via specs.set_dynamic_width or specs.set_dynamic_precision, and extracts the optional argument ID. At format time, the runtime value is retrieved from the argument pack and applied to the formatting operation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →