# How Rich's Syntax Class Implements Language-Specific Syntax Highlighting

> Discover how Rich's Syntax class achieves language-specific syntax highlighting. Learn about its Pygments integration, style mapping, and terminal text rendering.

- Repository: [Textualize/rich](https://github.com/Textualize/rich)
- Tags: internals
- Published: 2026-03-06

---

**Rich's `Syntax` class implements language-specific syntax highlighting by leveraging Pygments for lexical analysis and tokenization, then mapping those tokens to Rich `Style` objects through a pluggable theme system, ultimately building styled `Text` instances that render in the terminal.**

The `Syntax` class in the [Textualize/Rich](https://github.com/Textualize/rich) repository provides a flexible, high-performance pipeline for syntax highlighting that supports any programming language recognized by Pygments. By decomposing the process into three distinct stages—lexer resolution, theme mapping, and token styling—the class delivers customizable, colorful code display without spawning external processes.

## Lexer Selection: Determining the Language Parser

When instantiating a `Syntax` object, the class must first identify which Pygments lexer to use for parsing the source code. This logic resides in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py) and handles three distinct input scenarios.

### Resolving String Aliases to Pygments Lexers

If the `lexer` argument provided to `Syntax.__init__` is a string, the `lexer` property lazily resolves it via Pygments' `get_lexer_by_name`:

```python
if isinstance(self._lexer, Lexer):
    return self._lexer
try:
    return get_lexer_by_name(self._lexer, stripnl=False, ensurenl=True, tabsize=self.tab_size)
except ClassNotFound:
    return None  # Fallback handled later

```

This call attempts to match the string alias (e.g., `"python"`, `"javascript"`) to a registered Pygments lexer. If successful, it returns a configured lexer instance with newline and tab handling options set according to the `Syntax` instance's configuration.

### Guessing the Lexer from Filename and Content

When no lexer is specified, the `guess_lexer` class method attempts inference using `guess_lexer_for_filename`, which analyzes both the file extension and content heuristics. If this fails, it falls back to extension-based lookup. The `default_lexer` property provides a final safety net:

```python
return get_lexer_by_name("text", stripnl=False, ensurenl=True, tabsize=self.tab_size)

```

This ensures that even completely unknown file types render as plain text rather than raising exceptions.

## Theme Resolution: Mapping Tokens to Styles

Once the lexer is determined, the `Syntax` class must resolve how Pygments token types translate to Rich's internal `Style` objects. This abstraction is handled by the `SyntaxTheme` ABC and its concrete implementations in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py).

### The SyntaxTheme Abstraction

The `Syntax.get_theme` method determines which theme implementation to instantiate. If the `theme` argument matches a key in the built-in `RICH_SYNTAX_THEMES` dictionary (`"ansi_light"` or `"ansi_dark"`), it creates an `ANSISyntaxTheme`; otherwise, it instantiates a `PygmentsSyntaxTheme` with the supplied name (defaulting to `"monokai"`):

```python
if name in RICH_SYNTAX_THEMES:
    theme = ANSISyntaxTheme(RICH_SYNTAX_THEMES[name])
else:
    theme = PygmentsSyntaxTheme(name)

```

### PygmentsSyntaxTheme: Wrapping Pygments Styles

The `PygmentsSyntaxTheme` class wraps any Pygments style class. For each token type encountered during highlighting, it calls `style_for_token` and constructs a corresponding Rich `Style` object containing color, background color, and attributes (bold, italic, underline). Results are cached in `_style_cache` to avoid recomputing styles for repeated token types.

### ANSISyntaxTheme: Static ANSI Color Maps

For terminals with limited color support, `ANSISyntaxTheme` uses static mappings from Pygments token tuples to Rich `Style` objects defined in the `ANSI_LIGHT` and `ANSI_DARK` dictionaries. The lookup walks the token hierarchy (e.g., `Token.Keyword` → `Token`) until finding a matching style, then caches the result for subsequent tokens.

## The Highlighting Pipeline: From Tokens to Styled Text

The core transformation occurs in the `highlight` method, which orchestrates the conversion of source code into a Rich `Text` object ready for terminal rendering.

### Token Processing with `Syntax.highlight`

After preprocessing the source via `_process_code`, the method obtains the lexer (falling back to `default_lexer` if necessary) and iterates over the token stream:

```python
text.append_tokens(
    (token, _get_theme_style(token_type))
    for token_type, token in lexer.get_tokens(code)
)

```

Here, `lexer.get_tokens(code)` yields `(TokenType, str)` pairs where *TokenType* is a tuple like `('Token', 'Keyword')`. The `_get_theme_style` callable references the theme's `get_style_for_token` method, returning a Rich `Style` for each token. The `Text.append_tokens` method efficiently builds a single `Text` instance whose internal spans carry the style information.

### Handling Line Ranges for Performance

When highlighting only a subset of lines (via the `line_range` parameter), the class uses `line_tokenize` followed by `tokens_to_spans` to limit tokenization to the requested lines. This avoids processing entire large files when only a specific range is needed, significantly reducing memory consumption and CPU usage.

### Applying Custom Stylized Ranges

After the initial token pass, any user-defined stylized ranges added via `stylize_range` are applied through `_apply_stylized_ranges`. This method converts line/column positions into string offsets using `_get_code_index_for_syntax_position`, then applies additional styles via `Text.stylize` or `Text.stylize_before` to highlight specific variables, functions, or error lines independently of the lexer tokens.

Finally, the constructed `Text` object is combined with line numbers, indent guides, and background processing in `__rich_console__` and `_get_syntax` for rendering.

## Practical Code Examples

### Basic Python Highlighting with Pygments Theme

```python
from rich.console import Console
from rich.syntax import Syntax

code = '''
def hello(name):
    print(f"Hello, {name}!")
'''

syntax = Syntax(code, "python", theme="monokai", line_numbers=True)
Console().print(syntax)

```

This uses the default Pygments Python lexer and the "monokai" theme through the `PygmentsSyntaxTheme` implementation.

### Rendering Files with ANSI Dark Theme

```python
from rich.console import Console
from rich.syntax import Syntax

syntax = Syntax.from_path(
    "example.js",
    lexer="javascript",
    theme="ansi_dark",
    line_numbers=True,
    highlight_lines={3, 5},
)
Console().print(syntax)

```

Specifying `theme="ansi_dark"` triggers the `ANSISyntaxTheme` class using the `ANSI_DARK` color map, ensuring compatibility with basic terminal color support.

### Custom Token-to-Style Mapping

```python
from rich.console import Console
from rich.syntax import Syntax, ANSISyntaxTheme
from rich.style import Style
from pygments.token import Token

my_map = {
    Token.Keyword: Style(color="magenta", bold=True),
    Token.Name.Function: Style(color="cyan"),
    Token.Literal.String: Style(color="green"),
}
custom_theme = ANSISyntaxTheme(my_map)

syntax = Syntax(
    code='print("hello")',
    lexer="python",
    theme=custom_theme,
    line_numbers=False,
)
Console().print(syntax)

```

This bypasses string-based theme lookup by passing a `SyntaxTheme` instance directly, allowing fine-grained control over specific token styling.

## Summary

- **Rich's `Syntax` class** resides in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py) and uses Pygments lexers via `get_lexer_by_name` and `guess_lexer` to parse code into token streams, supporting any language Pygments recognizes.
- **Theme abstraction** is provided by the `SyntaxTheme` ABC, with `PygmentsSyntaxTheme` wrapping full Pygments styles and `ANSISyntaxTheme` providing ANSI-safe static color mappings for limited terminals.
- **The `highlight` method** converts lexer tokens into Rich `Text` objects by mapping token types to `Style` instances, utilizing `_style_cache` for performance and supporting `line_range` extraction to optimize large file handling.
- **Custom ranges** can be applied via `stylize_range`, which uses `_get_code_index_for_syntax_position` to convert line/column coordinates into string offsets for targeted styling.

## Frequently Asked Questions

### How does Rich's Syntax class handle unsupported programming languages?

If `get_lexer_by_name` cannot locate the specified lexer and `guess_lexer` fails to infer a language from the filename or content, the `default_lexer` property in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py) returns a plain-text lexer. This ensures the code displays without errors, though without language-specific token coloring.

### What is the difference between PygmentsSyntaxTheme and ANSISyntaxTheme?

`PygmentsSyntaxTheme` dynamically converts any Pygments style class (like "monokai" or "dracula") into Rich `Style` objects by calling Pygments' `style_for_token` method. `ANSISyntaxTheme` uses predefined static dictionaries (`ANSI_LIGHT` and `ANSI_DARK`) that map Pygments token types directly to terminal-safe ANSI colors, guaranteeing compatibility even on terminals with limited color support.

### Can I use Rich's Syntax class without Pygments installed?

No, Pygments is a required dependency for the `Syntax` class. The implementation imports from `pygments.lexers` and `pygments.token` to perform lexical analysis. Without Pygments, the lexer resolution in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py) would fail to generate the token streams necessary for the highlighting pipeline.

### How does Rich optimize performance when highlighting large files?

The `Syntax` class implements several optimizations in [`rich/syntax.py`](https://github.com/Textualize/rich/blob/main/rich/syntax.py): it maintains a `_style_cache` dictionary to avoid recomputing `Style` objects for repeated token types, and when a `line_range` parameter is provided, it uses `line_tokenize` and `tokens_to_spans` to process only the requested lines rather than tokenizing the entire file, significantly reducing memory overhead and processing time for partial views.