How bat Handles Non-Printable Characters: A Deep Dive into the Source Code

bat visualizes non-printable characters through a configurable replacement system that converts control codes, tabs, spaces, and newlines into readable symbols when the --show-all flag is enabled.

The bat command-line tool enhances standard cat output with syntax highlighting and Git integration, but its handling of non-printable characters requires special consideration for terminal safety. When displaying files containing binary data, control sequences, or whitespace characters, bat employs a sophisticated preprocessing pipeline defined in its Rust source code to ensure every byte is rendered visibly without breaking the terminal.

The Architecture Behind Non-Printable Character Handling

The replacement logic centers on three tightly-coupled components in the bat codebase: configuration flags, notation selection, and the preprocessor implementation.

Configuration Flags in Config::show_nonprintable

The primary gatekeeper is the show_nonprintable boolean flag defined in src/config.rs. This flag is populated by the command-line argument --show-all (or -A) in src/bin/bat/app.rs at line 387. When enabled, this flag triggers the replacement pipeline for every line processed by the printer.

Notation Selection with Config::nonprintable_notation

Users can choose between two visualization schemes via the nonprintable_notation enum defined in src/nonprintable_notation.rs:

  • Caret notation: Displays control characters as ^@, ^A through ^Z, ^[, ^\, ^], ^^, ^_, and ^? for DEL
  • Unicode notation: Uses Unicode Control Pictures block (U+2400–U+243F) such as ␀, ␁, ␊, ␡

The replace_nonprintable Preprocessor

The core transformation logic resides in replace_nonprintable within src/preprocessor.rs (lines 59–119). This function accepts a raw byte slice, tab width configuration, and notation preference, returning a String where every non-printable byte has been substituted with its visual representation.

How bat Processes Non-Printable Characters Step-by-Step

When bat prints a file with --show-all enabled, the following pipeline executes for every line:

  1. Line Acquisition: Printer::print_line in src/printer.rs receives a &[u8] buffer containing the raw file bytes.

  2. Conditional Replacement: At lines 146–155 in src/printer.rs, the code checks self.config.show_nonprintable. If true, it invokes replace_nonprintable with the line buffer, tab width, and notation settings; otherwise, it writes the buffer directly.

    if self.config.show_nonprintable {
        let line = replace_nonprintable(
            line_buffer,
            self.config.tab_width,
            self.config.nonprintable_notation,
        );
        write!(handle, "{line}")?;
    } else {
        // direct write
    }
  3. Byte-by-Byte Processing: Inside replace_nonprintable, the function iterates through the buffer, attempting UTF-8 decoding via try_parse_utf8_char. Based on the resulting character or byte value, it applies specific substitutions.

  4. Character Classification and Substitution:

    • Space (' '): Replaced with · (middle dot)
    • Tab ('\t'): Expanded to a visual tab stop using box-drawing characters (├──┤) respecting the configured tab_width
    • Line feed ('\n'): Replaced with caret (^J) or Unicode (␊) notation, followed by the actual newline
    • Control codes (0x00–0x1F): Mapped to caret notation (^@ through ^_) or Unicode Control Pictures (␀–␟)
    • Delete (0x7F): Rendered as ^? or ␡
    • Printable ASCII: Emitted unchanged
    • Other bytes: Escaped using Rust's escape_unicode format (e.g., \u{00E4})
  5. Output Generation: The function returns a complete String containing the visualized line, which Printer::print_line writes to the output handle.

Character Mapping Reference

The following table details how bat visualizes each character class when --show-all is enabled:

Character Type Byte Value Caret Notation Unicode Notation
Null 0x00 ^@ ␀
Control A–Z 0x01–0x1A ^A–^Z ␁–␚
Escape 0x1B ^[ ␛
File Separator 0x1C ^\ ␜
Group Separator 0x1D ^] ␝
Record Separator 0x1E ^^ ␞
Unit Separator 0x1F ^_ ␟
Space 0x20 · ·
Tab 0x09 Box-drawing Box-drawing
Line Feed 0x0A ^J + LF ␊ + LF
Delete 0x7F ^? ␡
Non-ASCII 0x80+ \u{XXXX} \u{XXXX}

Practical Examples

Command-Line Usage

To visualize all non-printable characters in a file using the default Unicode notation:

bat -A example.txt

To use caret notation instead:

bat --show-all --nonprintable-notation=caret binary-file.dat

Programmatic Usage in Rust

You can leverage bat's preprocessor directly in Rust code:

use bat::preprocessor::{replace_nonprintable, NonprintableNotation};

fn main() {
    let raw_bytes = b"\tHello\x00World\n";
    
    let visualized = replace_nonprintable(
        raw_bytes,
        4,                                    // tab width
        NonprintableNotation::Caret,         // or NonprintableNotation::Unicode
    );
    
    println!("{}", visualized);
}

Output (caret notation):


├──┤Hello^@World^J

Output (unicode notation):


├──┤Hello␀World␊

Summary

  • bat does not filter non-printable characters; it visualizes them deterministically when --show-all (-A) is enabled.
  • Configuration is controlled by Config::show_nonprintable and Config::nonprintable_notation in src/config.rs.
  • Two notation systems are available: caret notation (^M) and Unicode Control Pictures (␍), defined in src/nonprintable_notation.rs.
  • Core logic resides in replace_nonprintable within src/preprocessor.rs, which handles byte-by-byte substitution, tab expansion, and UTF-8 decoding.
  • Integration occurs in src/printer.rs where Printer::print_line conditionally invokes the preprocessor based on configuration.

Frequently Asked Questions

How do I show all non-printable characters in bat?

Use the --show-all flag (or -A for short). This enables the replacement logic that visualizes spaces as ·, tabs as box-drawing characters, and control codes using either caret or Unicode notation depending on your configuration.

What is the difference between caret and Unicode notation in bat?

Caret notation represents control characters using the ^ symbol followed by a letter or punctuation mark (e.g., ^G for bell, ^J for line feed). Unicode notation uses the Unicode Control Pictures block (e.g., ␇ for bell, ␊ for line feed). You can select between them using the --nonprintable-notation flag.

Does bat modify the actual file content when showing non-printable characters?

No, bat only modifies the displayed output. The replace_nonprintable function in src/preprocessor.rs creates a new visual representation of the data for terminal output, leaving the original file bytes untouched. This is a read-only transformation that occurs during the printing pipeline in src/printer.rs.

How does bat handle UTF-8 characters that are not ASCII?

For bytes outside the ASCII printable range (0x20–0x7E) that form valid UTF-8 sequences, bat attempts to decode them using try_parse_utf8_char within the preprocessor. If decoding succeeds, the character is emitted unchanged. If the bytes are invalid UTF-8 or represent non-printable control codes, they are escaped using Rust's escape_unicode format (e.g., \u{00E4}) or replaced with the configured notation symbols.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →