How fmtlib Calculates Unicode Display Width for CJK Characters and Emoji

fmtlib determines the visual column width of any Unicode code point by querying a compile-time lookup table of East Asian Width properties, returning 2 for wide characters (including CJK and emoji) and 1 for all others.

The {fmt} library (fmtlib/fmt) provides precise column alignment for international text by calculating display widths at formatting time. When you use width specifiers like {:>10} with strings containing Chinese, Japanese, or Korean characters, the library accounts for the fact that these glyphs occupy two terminal columns rather than one.

The Core Algorithm: display_width_of

The heart of fmtlib’s Unicode display width calculation resides in include/fmt/format.h within the function fmt::detail::display_width_of. This constexpr function takes a uint32_t code point and returns the number of display columns (either 1 or 2) that the character consumes in a monospace terminal.

Compile-Time Wide Character Table

Rather than relying on external libraries or runtime OS calls, fmtlib embeds a static table named wide_cp_data::ranges directly in the header. This table is generated from Unicode 16.0.0 TR-11 and contains sorted intervals of code points whose East_Asian_Width property is classified as W (Wide) or F (Fullwidth).

The table includes:

  • Hangul Jamo and Syllables
  • CJK Unified Ideographs and Radicals
  • Hiragana and Katakana
  • Full-width punctuation and symbols
  • Various emoji blocks

You can find the range definitions starting around line 562 in include/fmt/format.h. Because this data is constexpr, the compiler bakes the entire lookup structure into the binary with zero runtime initialization overhead.

Binary Search Implementation

The display_width_of function implements an optimized binary search over the ranges array. It first applies a fast path: if cp < 0x1100, the function immediately returns 1, as all code points below this threshold (primarily Basic Latin and extended Latin) are guaranteed narrow.

For code points at or above 0x1100, the algorithm performs a standard binary search (lines 820-834):

FMT_CONSTEXPR inline auto display_width_of(uint32_t cp) noexcept -> size_t {
  if (cp < 0x1100) return 1;
  size_t lo = 0;
  size_t hi = sizeof(wide_cp_data<>::ranges) / sizeof(wide_cp_range);
  while (lo < hi) {
    size_t mid = lo + (hi - lo) / 2;
    if (cp < wide_cp_data<>::ranges[mid].first)
      hi = mid;
    else if (cp > wide_cp_data<>::ranges[mid].last)
      lo = mid + 1;
    else
      return 2;               // inside a wide interval → width 2
  }
  return 1;                   // not wide → width 1
}

This approach provides O(log n) lookup time where n is the number of wide character ranges (typically a few dozen comparisons), making it suitable for high-performance formatting loops.

Special Handling for Emoji and Full-Width Blocks

Certain emoji blocks receive explicit full-width treatment regardless of their official East_Asian_Width classification. According to the source code in include/fmt/format.h (lines 785-808), the ranges table includes intervals for:

  • Miscellaneous Symbols and Pictographs
  • Supplemental Symbols and Pictographs
  • Emoticons and Transport/Map symbols

This ensures that modern emoji render correctly with double-column padding in terminal environments, even when the Unicode standard assigns them a neutral width property.

Integration with Format String Width Specifiers

When you specify a width in a format string (e.g., {:<width} or {:^width}), fmtlib calculates the total display width of the rendered string by summing the results of display_width_of for each UTF-8 decoded code point. The formatting logic (around lines 1895-1910 in include/fmt/format.h) then computes the necessary padding:

size_t width = compute_display_width(formatted_string);
unsigned spec_width = to_unsigned(specs.width);
size_t padding = spec_width > width ? spec_width - width : 0;
// padding is then written using the chosen fill character

This mechanism ensures that a CJK character counts as two columns when calculating how many fill characters (spaces or otherwise) to insert for right, left, or center alignment.

Practical Formatting Examples

The following program demonstrates how fmtlib respects double-width characters during padding operations:

#include <fmt/core.h>

int main() {
    // Simple ASCII – width 1 per character
    fmt::print("{:>10}\n", "hello");               // prints 5 spaces + "hello"

    // CJK string – each character counts as width 2
    fmt::print("{:>10}\n", u8"你好");               // prints 6 spaces + "你好"
    // The two Chinese characters occupy 4 columns, so only 6 spaces are added.

    // Mixed ASCII and CJK
    fmt::print("{:>10}\n", u8"abc你好");            // "abc" = 3, "你好" = 4 → total 7
    // => 3 spaces of padding to reach width 10.
}

The output remains perfectly column-aligned because fmtlib correctly identifies that "你好" consumes four display columns, not two.

Summary

  • display_width_of in include/fmt/format.h serves as the single source of truth for character widths, returning 1 or 2 based on Unicode East Asian Width properties.
  • wide_cp_data::ranges provides a compile-time table of wide character intervals derived from Unicode 16.0.0, enabling zero-overhead lookups.
  • Binary search over the sorted range table guarantees efficient O(log n) performance for every code point processed.
  • Emoji blocks are explicitly included in the wide table to ensure proper terminal rendering regardless of Unicode classification nuances.
  • Width specifiers automatically invoke this calculation when padding strings, ensuring accurate alignment for mixed ASCII/CJK content.

Frequently Asked Questions

How does fmtlib determine if a character is double-width?

fmtlib checks if the code point falls within any interval defined in the wide_cp_data::ranges table. This table contains all Unicode code points with East Asian Width properties W (Wide) or F (Fullwidth), plus specific emoji blocks. If the code point is inside a range, display_width_of returns 2; otherwise it returns 1.

Why does fmtlib use a binary search instead of a hash table or direct array?

The binary search approach minimizes binary size and memory usage. A direct array indexed by code point would require over a megabyte of storage (0x10FFFF entries), while a hash table would incur collision resolution overhead. The sorted interval table contains only a few hundred ranges, making binary search extremely cache-friendly and constexpr-evaluable at compile time for string literals.

Which Unicode version does fmtlib use for East Asian Width data?

The wide_cp_data table is generated from Unicode 16.0.0 TR-11 (East Asian Width specification). This includes modern CJK extensions, historic scripts, and recent emoji additions, ensuring accurate width calculation for contemporary text.

How does this affect string padding and alignment?

When you use width specifiers like {:>10}, fmtlib calculates the total display width by summing individual character widths. If you format the string "你好" (two CJK characters) with {:>10}, the library recognizes it already occupies 8 columns (2 characters × 2 columns each plus internal spacing logic), inserting only 2 spaces of padding rather than 8, keeping the visual output aligned with ASCII strings of the same format specification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →