How Rich Handles ANSI Escape Codes and Terminal Control Sequences

Rich processes ANSI escape codes through a three-stage pipeline that tokenizes raw strings with regular expressions, decodes them into structured Text and Style objects via AnsiDecoder, and renders platform-agnostic output while silently discarding unsupported control sequences.

The Textualize/rich library implements a robust parsing layer for ANSI escape codes that bridges raw terminal control sequences and its internal object model. Understanding how Rich handles these sequences—spanning from simple color codes to complex OSC hyperlinks—requires examining the tokenizer, decoder, and platform-specific renderers that work together to produce consistent terminal output across operating systems.

Tokenization and Regular Expression Parsing

Rich’s ANSI handling begins with strict tokenization defined in rich/ansi.py. The module compiles a comprehensive regular expression, re_ansi (lines 10‑15), that captures three distinct groups of escape sequences: simple ESC prefixes, OSC (Operating System Command) strings bounded by \x1b]...\\x1b\\, and CSI/SGR sequences following the standard \x1b[... pattern. This regex ensures that all incoming control characters are identified, whether supported or not.

The _ansi_tokenize function (lines 28‑56) iterates over input strings using this regex, yielding _AnsiToken objects that separate plain text payloads from SGR (Select Graphic Rendition) and OSC metadata. Each token maintains the original sequence information, allowing downstream processors to selectively handle specific code types while preserving text boundaries.

The AnsiDecoder State Machine

Central to Rich’s ANSI support is the AnsiDecoder class (lines 20‑122 in rich/ansi.py), instantiated once per Console session (as seen in rich/console.py at line 304). This decoder maintains a mutable Style instance that tracks current text attributes across token boundaries.

When decode_line processes a token stream, it handles three distinct operations:

  • Plain text emission – Appends literal characters to a Text object using the active style.
  • OSC 8 hyperlink detection – Recognizes hyperlink sequences (lines 57‑62) and invokes Style.update_link to associate URLs with subsequent text spans.
  • SGR style resolution – Splits compound SGR codes (e.g., 1;34m) on semicolons and maps them via SGR_STYLE_MAP (lines 59‑117). Standard attributes like bold (1), italic (3), and underline (4) update the style directly, while reset codes (0) clear the state.

Color Handling and Extended Palettes

Rich’s decoder natively supports 8‑bit and 24‑bit color extensions without external dependencies. When encountering 38;5;<n> or 48;5;<n> sequences, the decoder invokes Color.from_ansi (lines 77‑85) to map the 256‑color palette. For true‑color output using 38;2;<r>;<g>;<b> syntax, Color.from_rgb (lines 86‑93) constructs precise RGB color objects. These colors attach to the current Style instance, ensuring that subsequent text inherits the correct foreground or background until a reset code or new SGR sequence appears.

Console Integration and Platform Abstraction

The Console class in rich/console.py embeds an AnsiDecoder instance (line 304) and routes all potentially contaminated input through decode_line before rendering. This integration means that console.print() automatically normalizes ANSI strings into Rich’s internal Text representation, enabling seamless mixing of pre‑styled subprocess output with native Rich markup.

For Windows environments where native ANSI support may be limited, Rich delegates to rich/_windows_renderer.py. This renderer translates the finalized Style objects into Windows Console API calls rather than emitting escape sequences, ensuring consistent visual output without requiring Windows Terminal or ANSI.SYS drivers. Similarly, rich/pager.py reuses the same AnsiDecoder to preserve styles when buffering paged output.

Practical Code Examples

from rich.console import Console

console = Console()

# Standard SGR styling: bold blue text

console.print("\x1b[1;34mBold blue text\x1b[0m followed by normal text")

# 24-bit true color

console.print("\x1b[38;2;255;100;0mTrue-color orange\x1b[0m")

# OSC 8 hyperlink

console.print("\x1b]8;;https://rich.readthedocs.io\x1b\\Rich Documentation\x1b]8;;\x1b\\")

Each call routes the string through AnsiDecoder.decode_line, which tokenizes the escape sequences, updates internal style states, and returns a Text object that the console renderer outputs according to the host terminal’s capabilities.

Summary

  • Tokenization: rich/ansi.py uses re_ansi (lines 10‑15) and _ansi_tokenize (lines 28‑56) to split input into plain text and control sequences.
  • Decoding: AnsiDecoder (lines 20‑122) maintains a mutable Style state, mapping SGR codes to style attributes and OSC 8 sequences to hyperlinks.
  • Color support: Extended 8‑bit and 24‑bit colors are parsed via Color.from_ansi and Color.from_rgb (lines 77‑93).
  • Integration: Console initializes AnsiDecoder at line 304 of rich/console.py, automatically normalizing ANSI input for the rendering pipeline.
  • Cross‑platform: Windows terminals receive fallback handling through rich/_windows_renderer.py, converting styles to native API calls when ANSI streams are unavailable.

Frequently Asked Questions

How does Rich handle unsupported or malformed ANSI sequences?

Rich silently ignores any escape sequences that do not match supported SGR or OSC patterns. During the tokenization phase in _ansi_tokenize, unrecognized CSI sequences are still captured as tokens but are discarded by AnsiDecoder during style resolution, ensuring that garbage sequences do not crash the renderer or corrupt output.

Does Rich support cursor movement and terminal clearing codes?

No. While the re_ansi regex recognizes cursor movement and other control sequences, AnsiDecoder specifically processes only styling-related SGR codes and OSC 8 hyperlinks. Cursor positioning commands (such as \x1b[2J or \x1b[H) are tokenized but ignored during the conversion to Text objects, as Rich’s rendering model manages layout independently of absolute cursor positioning.

Can Rich convert ANSI-styled strings into plain text without escape codes?

Yes. By passing a string containing ANSI codes through AnsiDecoder.decode_line, you obtain a Text object. Rendering this object through Console.print() with a legacy_windows or file output that does not support ANSI will produce plain text without escape sequences. Additionally, the Text object’s plain attribute contains the raw string content stripped of all formatting.

How does Rich ensure ANSI compatibility on legacy Windows systems?

On Windows platforms lacking native ANSI support, Rich utilizes rich/_windows_renderer.py to translate Style attributes into Windows Console API calls (SetConsoleTextAttribute). This bypasses the ANSI stream entirely, allowing color and style output on legacy cmd.exe windows while the AnsiDecoder still handles the initial parsing of any ANSI input strings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →