How Rich Implements Markup Parsing: Inside the [bold] and [color] Syntax Engine

Rich parses markup syntax like [bold] and [color] using a three-phase pipeline in rich/markup.py that tokenizes input with the RE_TAGS regular expression, builds a stream of Tag objects via _parse(), and renders them into styled text spans using a stack-based approach in render().

Rich is a popular Python library for rendering rich text and beautiful formatting in the terminal. According to the Textualize/rich source code, the library's markup parsing engine transforms human-readable tags into structured style spans through a sophisticated process defined in rich/markup.py.

The Three-Phase Parsing Pipeline

The core markup engine operates in three distinct phases: tokenization, parsing, and rendering. Each phase transforms the input string progressively until it becomes a Text object containing styled spans.

Phase 1: Tokenization with RE_TAGS

The process begins with a single regular expression that scans the input for markup tags. In rich/markup.py (lines 12-14), the RE_TAGS pattern is defined as:

RE_TAGS = re.compile(r"""((\\*)\[([a-z#/@][^[]*?)])""", re.VERBOSE)

This regex identifies three distinct groups: the full matched text, any preceding backslashes, and the tag body (tag_text). The parser detects tags like [name], [name=param], and handles escaped brackets by counting backslashes—odd counts indicate escaped brackets that should be treated as plain text, while even counts indicate valid tag openers.

Phase 2: Building the Stream with _parse()

The _parse() function (lines 73-104) walks the tokenized input and yields a stream of tuples containing:

  • Plain text chunks that appear between tags
  • Tag objects represented as Tag(name, parameters) namedtuples

Closing tags are expressed with a leading slash, such as [/] or [/bold]. The Tag namedtuple stores the tag name (e.g., bold, red, on blue) and optional parameters, providing a structured representation of the markup syntax for the rendering phase.

Phase 3: Rendering with the Style Stack via render()

The render() function (lines 106-131) drives the core markup engine by maintaining a style stack (style_stack: List[Tuple[int, Tag]]). When the renderer encounters an opening tag, it pushes the current text length and the normalized Tag onto the stack. Upon finding a closing tag, pop_style() removes the matching opening entry—or the most recent entry for an implicit [/].

For each text segment, the parser calls append() to add characters to a Text instance while accumulating style spans as Span(start, end, style) objects. Any tags left open at the end of processing are closed automatically, converting them into final Span objects. The resulting Text instance carries the plain string plus its style spans, ready for the console renderer to convert into ANSI escape sequences.

Style Resolution and Normalization

Opening tags undergo normalization via Style.normalize() (defined in rich/style.py), which converts shortcuts like [b] into "bold" and validates color names. The Tag class provides a markup property that generates helpful error messages when closing tags fail to match their corresponding openers.

For meta-tags such as [@handler], the render() function extracts the handler name and optional parameters, evaluates them using literal_eval, and stores them as meta-styles using Style(meta={…}). This allows the console to react to interactive elements like click handlers while maintaining the same parsing infrastructure.

Public API: Converting Markup to Text

Rich exposes the markup engine through the Text.from_markup() method in rich/text.py (lines 60-70). This public API simply calls the internal render() function and returns the populated Text object:

from rich.text import Text

markup = "[bold red]Hello [italic]World[/italic]![/]"
styled_text = Text.from_markup(markup)

The resulting styled_text contains two spans: characters 0-11 carry the style bold red, while characters 6-11 carry italic. When printed with a Console instance, Rich translates these spans into the appropriate terminal control sequences.

Practical Code Examples

Basic Bold and Color Markup

from rich.console import Console
from rich.text import Text

console = Console()
txt = Text.from_markup("[bold blue]Rich[/] makes it easy!")
console.print(txt)  # Displays "Rich" in bold blue, followed by normal text

Nested Tags with Explicit Closing

txt = Text.from_markup("[red]Error: [bold]File not found[/bold][/red]")

# Outer red tag spans the entire string

# Inner bold tag applies only to "File not found"

Hexadecimal Color Values

txt = Text.from_markup("[#ff00ff]Magenta text[/]")

# The hex color is parsed by the Color class and stored in the span's style

Meta Handler Tags

txt = Text.from_markup("[@click=handle_click]Click me[/]")

# The @click tag is stored as a meta-style for later console interaction

Key Implementation Files

File Responsibility
rich/markup.py Core markup parser containing RE_TAGS, _parse(), and render()
rich/style.py Style object definition, color name parsing, and attribute normalization
rich/text.py Public API (Text.from_markup()) that constructs Text from markup strings
rich/console.py Consumes Text objects and emits final ANSI escape sequences

Summary

  • Rich markup parsing operates through a three-phase pipeline in rich/markup.py: tokenization with RE_TAGS, stream building with _parse(), and rendering with render().
  • The style stack mechanism tracks nested tags by pushing Tag objects at opening tags and popping them at closing tags, generating Span objects that define style ranges.
  • Style normalization occurs via Style.normalize() in rich/style.py, converting shortcuts like [b] to "bold" and validating color specifications.
  • Meta-tags starting with @ are handled specially, storing handler information as metadata within the style object.
  • The public entry point Text.from_markup() in rich/text.py provides a simple interface to the entire parsing engine.

Frequently Asked Questions

How does Rich handle escaped brackets in markup?

Rich detects escaped brackets by analyzing the backslash count in the RE_TAGS regex match. When the parser encounters an odd number of backslashes preceding a bracket, it treats the sequence as literal text rather than a tag opener, allowing users to display literal [ characters by writing \[ or \\[ depending on Python string escaping rules.

What is the difference between implicit and explicit closing tags?

Explicit closing tags like [/bold] target specific named styles, while implicit closing tags using [/] automatically close the most recently opened tag on the style stack. The render() function's pop_style() method handles both cases, matching explicit closers by name or defaulting to the last-in-first-out (LIFO) order for implicit closers.

How are meta tags like [@click] processed differently from style tags?

Meta tags starting with @ bypass standard style normalization and instead extract handler names and parameters using literal_eval. These are stored as Style(meta={handler: params}) objects rather than traditional color or attribute styles, enabling the console to attach interactive behaviors to text spans while using the same parsing infrastructure.

Where does the actual ANSI escape sequence generation happen?

While rich/markup.py produces Text objects containing style spans, the translation to ANSI escape sequences occurs in rich/console.py. The Console class interprets the Span objects and their associated styles, converting them into terminal-specific control codes during the final output phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →