# How Ghostty's Terminal Parser Uses SIMD Instructions for High-Performance UTF-8 Decoding

> Discover how Ghostty's terminal parser uses SIMD instructions for lightning-fast UTF-8 decoding. See how it processes data in 4KB chunks to eliminate branches and boost performance.

- Repository: [Ghostty/ghostty](https://github.com/ghostty-org/ghostty)
- Tags: internals
- Published: 2026-05-01

---

**Ghostty leverages SIMD instructions by routing printable UTF-8 text through vectorized decoding routines from the simdutf library while reserving scalar processing for control sequences, eliminating per-byte branches and processing data in 4KB chunks.**

The Ghostty terminal emulator (ghostty-org/ghostty) achieves low-latency text rendering by processing terminal output in two distinct stages. While traditional terminal emulators parse input byte-by-byte, Ghostty's terminal parser utilizes SIMD instructions to decode large batches of printable Unicode characters in parallel, falling back to a scalar state machine only when encountering escape sequences or malformed data.

## The Two-Stage Parsing Architecture

Ghostty splits terminal stream processing between a **scalar state machine** and a **SIMD fast path**. This separation allows the parser to handle complex VT sequences correctly while maximizing throughput for the common case of printable text.

### Scalar State Machine

The scalar path handles control sequences, state transitions, and malformed input byte-by-byte. Implemented primarily in `src/terminal/stream.zig`, the `next` function processes individual bytes through a complete VT220-compatible state machine. This path is necessary because escape sequences require contextual parsing that cannot be easily vectorized.

### SIMD Fast Path

When the parser is in the **ground state** (expecting only printable characters), the `nextSlice` method routes input to `nextSliceCapped`. This function processes data in chunks, calling vectorized UTF-8 decoders that operate on multiple bytes simultaneously. The SIMD path is only invoked when the parser state confirms no pending escape sequence, enabling zero-branch execution for the bulk of terminal output.

## Compile-Time SIMD Configuration

SIMD support is controlled at compile time via the `simd` build option defined in `src/build/Config.zig` and exposed through `src/terminal/build_options.zig`. When building with `zig build -Dsimd=false` or when `debug = true`, the compiler excludes the vectorized code paths and falls back to a pure-Zig scalar loop.

The build system integrates the **simdutf** library through `src/build/SharedDeps.zig`, linking C++ source files from `src/simd/*.cpp` that provide platform-specific SIMD implementations using SSE, AVX2, AVX-512, NEON, or VSX instructions depending on the target architecture.

## Chunked Processing and Buffer Management

The SIMD-enabled path processes input in **4KB chunks** that fit into a temporary buffer `cp_buf` of `u32` code-points. In `src/terminal/stream.zig`, the `nextSliceCapped` function:

1. Splits input into chunks capped at 4096 bytes
2. Allocates a stack buffer of 4096 `u32` values for decoded code-points
3. Calls the SIMD decoder via `simd.vt.utf8DecodeUntilControlSeq`

This chunking strategy balances cache efficiency with amortized function call overhead, ensuring the vectorized routines process enough data to outweigh the setup cost.

## Vectorized UTF-8 Decoding with simdutf

The actual SIMD implementation resides in `src/simd/vt.zig`, which exposes an `extern "c"` wrapper around the C++ function `ghostty_simd_decode_utf8_until_control_seq` from the simdutf library.

This routine performs three critical operations in a tight vectorized loop:

- **Vectorized scanning** for the ESC byte (`0x1B`) using SIMD comparison instructions
- **UTF-8 validation and decoding** of multi-byte sequences into Unicode code-points
- **Boundary detection** for incomplete sequences at chunk edges

The function returns the number of consumed input bytes and produced code-points, allowing the Zig parser to advance its position precisely.

## Zero-Branch Execution in Ground State

The SIMD path achieves its performance advantage through **zero-branch execution**. Because the parser invokes the vectorized routine only when in the ground state (no pending escape sequence), it avoids the per-byte conditional branch that would normally distinguish printable characters from control codes.

In a traditional byte-by-byte parser, the CPU must predict whether each byte is text or a control sequence, causing costly branch mispredictions on mixed content. Ghostty's approach eliminates this speculation by processing entire chunks of printable text in a single vectorized pass, then handling the control sequence separately when the SIMD routine hits an ESC byte or returns to the scalar decoder.

## Scalar Fallback and Edge Case Handling

When the SIMD decoder encounters incomplete UTF-8 sequences, malformed bytes, or the ESC character, control returns to the scalar path. The remaining unprocessed bytes are fed individually to `nextUtf8`, a function that handles:

- **Split escape sequences** spanning chunk boundaries
- **Invalid UTF-8** sequences requiring state machine error handling
- **Control characters** (bytes 0x00-0x1F and 0x7F) embedded in text

This hybrid approach guarantees correct terminal emulation while retaining the performance benefits of SIMD for well-formed, text-heavy output.

## SIMD in Kitty Graphics Decoding

Ghostty extends its SIMD usage beyond UTF-8 parsing to base64 decoding for Kitty graphics commands. In `src/terminal/kitty/graphics_command.zig`, the parser calls `simd.base64.decode` and `simd.base64.maxLen` to process image data efficiently.

These functions follow the same architectural pattern: Zig wrappers calling into C++ SIMD routines from the simdutf library, processing large blocks of data with vectorized instructions while maintaining a scalar fallback for edge cases.

## Summary

- Ghostty uses a hybrid architecture separating scalar control sequence handling from SIMD UTF-8 bulk decoding.
- The `simd` build option in `src/build/Config.zig` controls whether vectorized code paths are compiled.
- Input is processed in 4KB chunks by `nextSliceCapped` in `src/terminal/stream.zig`, which calls `simd.vt.utf8DecodeUntilControlSeq` from `src/simd/vt.zig`.
- The underlying C++ function `ghostty_simd_decode_utf8_until_control_seq` from the simdutf library scans for ESC bytes (`0x1B`) using platform-specific SIMD instructions.
- When the parser is in the ground state, it executes the SIMD path with zero branches for printable characters, eliminating costly branch mispredictions.
- Edge cases including malformed UTF-8, incomplete sequences, and split escape sequences fall back to the scalar `nextUtf8` decoder.

## Frequently Asked Questions

### What SIMD instruction sets does Ghostty's terminal parser use?

Ghostty delegates to the simdutf library, which automatically selects the optimal instruction set for the target platform, including SSE, AVX2, AVX-512, ARM NEON, or POWER VSX. The Zig code in `src/simd/vt.zig` wraps these C++ routines via `extern "c"` declarations, allowing the parser to call `ghostty_simd_decode_utf8_until_control_seq` without knowing the specific underlying architecture.

### How does the parser handle control sequences when using SIMD acceleration?

The SIMD decoder scans for the ESC byte (`0x1B`) using vectorized comparisons. When it encounters an escape sequence or incomplete UTF-8, it stops and returns the consumed byte count. The remaining input is processed by the scalar state machine in `src/terminal/stream.zig` via the `nextUtf8` function, ensuring correct VT sequence interpretation while maximizing throughput for printable characters.

### Can I disable SIMD optimizations for debugging purposes?

Yes. Set the `simd` build option to false using `zig build -Dsimd=false` or enable debug mode, which automatically disables SIMD. This compiles a pure-Zig scalar fallback that processes input sequentially through the `next` function, making the behavior identical but significantly slower on large outputs.

### Does Ghostty use SIMD for operations other than UTF-8 parsing?

Yes. The parser also uses SIMD acceleration for base64 decoding in Kitty graphics commands. The file `src/terminal/kitty/graphics_command.zig` calls `simd.base64.decode` and `simd.base64.maxLen` to process image data efficiently, following the same pattern of wrapping C++ SIMD routines from the simdutf library.