# How fmtlib Parses Format Specifiers: A Deep Dive into {fmt}'s State-Machine Parser

> Discover how fmtlib parses format specifiers with its constexpr state machine. Understand the strict order of alignment, sign, width, precision, and type components in fmtlib's core.h for efficient formatting.

- Repository: [Hello World Foundation/fmt](https://github.com/fmtlib/fmt)
- Tags: deep-dive
- Published: 2026-09-08

---

**`{fmt}` parses format specifiers using a constexpr state machine in [`include/fmt/core.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/core.h), driven by the `parse_format_specs` function that processes alignment, sign, width, precision, and type components in strict order.**

The `{fmt}` library implements one of the most robust format specifier parsers in modern C++. Unlike legacy `printf` parsing, `{fmt}` performs **compile-time validation** where possible, using a carefully structured state machine to ensure specifiers are syntactically correct and semantically appropriate for the argument type. This article examines the actual implementation in the [fmtlib/fmt](https://github.com/fmtlib/fmt) repository to explain how format specifier parsing works under the hood.

## The Entry Point: `parse_format_specs`

All format specifier parsing begins in [`include/fmt/core.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/core.h) at line 55, where the `parse_format_specs` function is defined:

```cpp
auto parse_format_specs(const Char* begin, const Char* end, 
                        dynamic_format_specs<Char>& specs, ...)

```

This constexpr function iterates through characters between `{` and `}`, building a **`format_specs`** object that stores every parsed component. The function returns an iterator positioned at the closing `}` or the end of the specifier string.

If the format string contains an empty specifier `{}`, the parser returns immediately without processing. Otherwise, it initializes a state machine to track which components have been encountered.

## State Machine Architecture

The parser uses an internal **`state` enum** to enforce correct ordering of specifier components. The states progress through: `start` → `align` → `sign` → `hash` → `zero` → `width` → `precision` → `locale`.

The helper function `enter_state` (lines 70-77) validates transitions:

```cpp
// From include/fmt/core.h
auto enter_state = [&](state new_state) {
  if (int(state) >= int(new_state)) report_error("invalid format spec");
  state = new_state;
};

```

This strict ordering prevents malformed specifiers like `{:10<}` (width before alignment) from compiling or running.

## Component-by-Component Parsing

### Alignment and Fill Characters

Alignment tokens `<`, `>`, and `^` are detected at lines 94-101. The parser supports **multibyte fill characters**—if a character precedes an alignment token, it is extracted as the fill sequence:

```cpp
// Lines 71-88: Extracting fill character
if (*begin != '{' && *begin != '}') {
  auto c = *begin++;  // consume potential fill character
  if (begin != end && *begin == '<' || *begin == '>' || *begin == '^') {
    specs.set_fill(c);  // line 84
    // ... set alignment based on *begin
  }
}

```

This enables patterns like `{:*>10}` for right-aligned, asterisk-filled output.

### Sign, Alternative Form, and Zero-Padding

- **Sign** (`+`, `-`, space): Lines 102-109, restricted to signed integers and floating-point types via `in(arg_type, sint_set | float_set)`
- **Alternative form** (`#`): Lines 111-114, enabling `0x` prefixes for hex, etc.
- **Zero-padding** (`0`): Lines 115-124, treated as numeric alignment with fill `'0'`, rejected for non-numeric arguments

Each modifier validates the argument type before acceptance, preventing nonsensical combinations like `{:+}` on a string.

### Width and Precision

**Width** is parsed by `parse_width` (lines 130-132), handling both literal values and dynamic specifications (`{}` or `{n}`):

```cpp
specs.set_dynamic_width(arg_id, index);  // line 133

```

**Precision** follows a dot (`.`) at lines 136-144, using `parse_precision` with similar dynamic value support (`specs.set_dynamic_precision` at line 149).

### Locale Flag and Presentation Type

The **`L`** flag (lines 140-142) enables locale-aware formatting for numeric separators. Finally, the **presentation type** (lines 146-166) determines output format through `parse_presentation_type`, which validates against type-specific allowed sets:

| Type Character | Meaning | Allowed Argument Types |
|---------------|---------|----------------------|
| `d`, `i` | Decimal integer | Integral types |
| `x`, `X` | Hexadecimal | Integral types |
| `o` | Octal | Integral types |
| `b`, `B` | Binary | Integral types |
| `f`, `F` | Fixed floating-point | Floating-point types |
| `e`, `E` | Scientific notation | Floating-point types |
| `g`, `G` | General format | Floating-point types |
| `a`, `A` | Hexadecimal floating-point | Floating-point types |
| `c` | Character | Integral/character types |
| `s` | String | String types |
| `p` | Pointer | Pointer types |
| `?` | Debug format | Any type with debug support |

## The `format_specs` Storage Structure

Parsed data is stored in **`struct format_specs : basic_specs`** (line 831), which inherits core flags and adds:

- Fill character (multibyte support)
- Dynamic width/precision tracking
- Argument ID references for runtime values

This structure is consumed by formatter `write` overloads in [`include/fmt/format.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/format.h) to produce final output.

## Practical Examples

These examples exercise the complete specifier parser:

```cpp
#include <fmt/core.h>
#include <fmt/format.h>

int main() {
    int v = 42;
    double pi = 3.14159;
    
    // Fill, alignment, width
    fmt::print("{:*^12}\n", v);      // *****42*****
    
    // Sign, zero-fill, width
    fmt::print("{:+08}\n", v);       // +0000042
    
    // Alternative form, hex type
    fmt::print("{:#x}\n", v);        // 0x2a
    
    // Precision and floating-point type
    fmt::print("{:.2f}\n", pi);      // 3.14
    
    // Locale-aware formatting
    fmt::print("{:L}\n", 1234);      // 1,234 (locale-dependent)
    
    // Debug type for pairs
    fmt::print("{:?}\n", std::make_pair(1, 2));  // (1, 2)
}

```

## Key Source Files

Understanding format specifier parsing requires familiarity with these files in the [fmtlib/fmt](https://github.com/fmtlib/fmt) repository:

- **[`include/fmt/core.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/core.h)** — Core definitions including `format_specs`, `basic_specs`, and the complete `parse_format_specs` implementation
- **[`include/fmt/format.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/format.h)** — Entry point for formatting; forwards parsed specs to writer functions
- **[`include/fmt/format-inl.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/format-inl.h)** — Inline utilities that consume `format_specs` for actual output generation

## Summary

- **`parse_format_specs`** in [`include/fmt/core.h`](https://github.com/fmtlib/fmt/blob/main/include/fmt/core.h) is the central parsing function, operating at compile-time where possible
- A **strict state machine** enforces correct specifier component ordering through the `state` enum and `enter_state` helper
- **Type validation** occurs at every stage, preventing invalid combinations like sign flags on strings
- **Dynamic specifiers** (`{}` for width/precision) are tracked separately for runtime resolution
- Results are stored in **`format_specs`** objects consumed by writer functions in the format implementation

## Frequently Asked Questions

### Why does `{fmt}` use a state machine for parsing format specifiers?

The state machine design enforces that format specifier components appear in a strict, unambiguous order. This prevents malformed patterns like `{:10<}` where width precedes alignment, and enables clear error messaging. The `enter_state` helper at lines 70-77 validates transitions by comparing enum values, catching errors at compile time for literal format strings.

### Can I use multibyte Unicode characters as fill characters in {fmt}?

Yes. The parser at lines 71-88 handles fill character extraction before checking for alignment tokens. The `specs.set_fill` call stores the complete character sequence, supporting Unicode fill characters when using `wchar_t` or `char8_t/char16_t/char32_t` specializations.

### How does {fmt} validate that a type character matches the argument type?

Each presentation type case calls `parse_presentation_type` with an allowed type set (lines 146-166). For example, `case 'x': return parse_presentation_type(pres::hex, integral_set)` restricts hex formatting to integral types. The function checks `in(arg_type, allowed_set)` and reports errors for mismatches, providing type-safe formatting that `printf` cannot match.

### What happens when width or precision is specified dynamically with `{}`?

The parser detects the opening brace in `parse_width` or `parse_precision`, records the dynamic nature via `specs.set_dynamic_width` or `specs.set_dynamic_precision`, and extracts the optional argument ID. At format time, the runtime value is retrieved from the argument pack and applied to the formatting operation.