How bat Handles Different File Encodings: UTF-8, UTF-16, and Binary Detection
bat natively supports UTF-8 and UTF-16 (LE/BE) through automatic detection via the content_inspector crate, decodes UTF-16 using encoding_rs with BOM removal, and treats all other encodings as UTF-8 with lossy fallback or binary suppression.
The sharkdp/bat repository implements a robust encoding detection and conversion system that allows the tool to display files with syntax highlighting across different character sets. Understanding how bat handles different file encodings reveals why it can seamlessly render UTF-16 formatted logs while maintaining UTF-8 as its default internal representation.
Encoding Detection Architecture in bat
When bat opens a file, it immediately inspects the first bytes to determine the content type before processing any further. This detection happens in src/input.rs through a coordinated sequence of operations that classify the input as UTF-8, UTF-16LE, UTF-16BE, or binary.
Content Type Inspection with content_inspector
The detection process begins in Input::open() at lines 60-66 of src/input.rs, where bat creates an InputReader and reads the first line of the file. Immediately after, at lines 64-69, bat calls content_inspector::inspect(&first_line) to analyze the byte patterns. This crate examines the initial bytes for UTF-16 Byte Order Marks (BOMs) and specific byte sequences that distinguish UTF-16LE, UTF-16BE, valid UTF-8, and binary content.
UTF-16 Line Reading Logic
If the inspection detects UTF-16LE or UTF-16BE, bat enters a specialized reading path. At lines 70-76 of src/input.rs, the code reads extra bytes until it locates a line terminator that aligns with UTF-16 byte boundaries. The InputReader::read_line() method at lines 90-98 routes the operation to read_utf16_line for UTF-16 content or to standard read_until('\n') for UTF-8, ensuring that multi-byte characters are not split across read operations.
How bat Decodes UTF-16 Files
Once bat has read the raw bytes, the decoding process shifts to src/printer.rs, where the tool converts UTF-16 byte sequences into displayable UTF-8 strings for terminal output.
BOM Removal and encoding_rs Integration
At lines 626-628 of src/printer.rs, bat utilizes the encoding_rs crate to handle UTF-16 decoding. The code calls decode_with_bom_removal() on the UTF-16 buffers, which automatically strips any Byte Order Mark from the output while converting the byte sequence to UTF-8. This ensures that files with or without BOMs display correctly without manual intervention.
Handling UTF-16LE and UTF-16BE Variants
The printer distinguishes between the two UTF-16 variants using explicit match arms. For ContentType::UTF_16LE, bat invokes UTF_16LE.decode_with_bom_removal(buf), and for ContentType::UTF_16BE, it uses UTF_16BE.decode_with_bom_removal(buf). This explicit handling ensures correct byte order interpretation regardless of platform endianness.
Binary File Handling and Fallback Behavior
For files that do not match UTF-8 or UTF-16 patterns, bat implements a fallback strategy that prioritizes terminal safety over forced display.
UTF-8 Lossy Conversion for Unknown Encodings
At lines 467-489 of src/printer.rs, bat treats unrecognized encodings as UTF-8. When encountering invalid byte sequences, the code uses String::from_utf8_lossy(), which replaces invalid UTF-8 sequences with the Unicode replacement character (). This allows bat to display partially corrupted files without crashing, though it cannot accurately render legacy encodings like ISO-8859-1 or Windows-1252.
Suppressing Binary Output vs. --show-all
When content_inspector classifies a file as ContentType::BINARY, bat suppresses output to prevent terminal corruption from binary data. By default, bat prints a warning indicating the file is binary and skips content display. However, passing the --show-all or -A flag disables this protection, forcing bat to process the file through the standard UTF-8 lossy path and display raw bytes.
Practical Examples: Working with Encodings in bat
The following commands demonstrate how bat handles different file encodings in practice:
Display a standard UTF-8 source file with syntax highlighting:
bat src/main.rs
View a UTF-16LE encoded log file automatically detected and decoded:
bat windows_log.txt
Attempt to display a binary file (suppressed by default):
bat image.png
# Output: "File: image.png <BINARY>"
Force display of binary content using --show-all:
bat -A image.png
Convert a legacy ISO-8859-1 file to UTF-8 before viewing with bat:
iconv -f ISO-8859-1 -t UTF-8 legacy.txt | bat -
Summary
- bat natively handles UTF-8 and UTF-16 (LE/BE) through automatic detection in
src/input.rsusing thecontent_inspectorcrate. - UTF-16 decoding occurs in
src/printer.rsviaencoding_rs, which removes BOMs and converts to UTF-8 for display. - Binary files are detected and suppressed by default to prevent terminal corruption, with
--show-all(-A) available to override. - Legacy encodings (ISO-8859-1, Windows-1252) are not supported; convert them to UTF-8 using external tools like
iconvbefore piping to bat.
Frequently Asked Questions
Does bat support ISO-8859-1 or Windows-1252 encodings?
No, bat does not natively support ISO-8859-1, Windows-1252, or other legacy single-byte encodings. When encountering these files, bat attempts to interpret them as UTF-8 and applies lossy conversion, replacing invalid sequences with replacement characters. To view these files correctly, convert them to UTF-8 first using iconv -f ISO-8859-1 -t UTF-8 file.txt | bat -.
How does bat detect UTF-16 versus UTF-8 automatically?
Bat uses the content_inspector crate to examine the first bytes of the input file immediately upon opening. In src/input.rs, the content_inspector::inspect() function analyzes byte patterns to identify UTF-16 Byte Order Marks and specific byte sequences that distinguish UTF-16LE, UTF-16BE, valid UTF-8, and binary content before any decoding occurs.
What happens when bat encounters a binary file?
When content_inspector classifies a file as binary, bat suppresses the output to prevent terminal corruption from control characters or escape sequences. By default, bat displays a warning message indicating the file is binary and skips content display. If you need to view the binary content, use the --show-all or -A flag, which forces bat to process the file through its standard output path.
Can I force bat to treat a file as UTF-16?
Bat does not provide a command-line flag to manually force UTF-16 interpretation. The tool relies entirely on automatic detection via content_inspector. If automatic detection fails for a UTF-16 file without a BOM, ensure the file has the proper byte order mark at the start, or convert the file to UTF-8 using external tools before viewing with bat.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →