How to Search for Non-UTF-8 Encoded Files with ripgrep: Latin-1, GBK, and More

Use the -E/--encoding flag to specify the source encoding (e.g., rg -E latin1 or rg -E gbk), which transcodes files to UTF-8 before searching, or use -E none to search raw bytes.

By default, the ripgrep command-line tool assumes every file is UTF-8 encoded. However, legacy systems and international projects often rely on alternative encodings like Latin-1 (ISO-8859-1), GBK, or Shift_JIS. According to the BurntSushi/ripgrep source code, you can search for non-UTF-8 encoded files by leveraging the built-in transcoding pipeline that converts source encodings to UTF-8 on the fly.

Understanding ripgrep’s Default Encoding Behavior

ripgrep operates on UTF-8 by default, but it includes a BOM (Byte Order Mark) sniffing mechanism. When it encounters a UTF-16 file with a BOM, it automatically handles transcoding regardless of command-line settings. This behavior is documented in GUIDE.md and implemented in src/core/main.rs, where the search logic first checks for UTF-16 BOM markers before applying user-specified encoding rules.

Using the -E/--encoding Flag to Search Non-UTF-8 Files

The -E or --encoding flag tells ripgrep to treat every file it searches as being in the specified encoding. According to crates/ignore/src/encoding.rs, ripgrep uses the encoding_rs library to map encoding names like latin1, gbk, euc-jp, or shift_jis to their respective character sets.

When you specify an encoding, ripgrep transcodes the file contents from that encoding to UTF-8 before feeding them to the regex engine. This allows Unicode-aware regex features—such as \w for word characters or . for any character—to work correctly on non-UTF-8 source files.

Searching Latin-1 (ISO-8859-1) Files

For Western European text files using single-byte Latin-1 encoding, specify the encoding explicitly:

rg "café" -E latin1 src/

ripgrep converts the Latin-1 bytes to UTF-8 internally, ensuring accented characters match correctly against your UTF-8 pattern.

Searching GBK and Other Multi-Byte Encodings

For Chinese text encoded in GBK, or Japanese text in Shift_JIS or EUC-JP, use the appropriate encoding label:


# Search for the string "错误" in GBK-encoded files

rg "错误" -E gbk path/to/file.txt

The transcoding layer handles the conversion from GBK to UTF-8, allowing the regex engine to process the content as valid Unicode.

Raw-Byte Searches with -E none

To bypass all encoding detection and transcoding—including BOM sniffing—use the special value none. This mode searches the file's raw bytes without conversion, which is useful for inspecting exact byte sequences or handling malformed files.


# Find exact UTF-16 byte sequences by searching raw bytes

rg '(?-u)\(\x045\x04@\x04;\x04>\x04:\x04' -E none -a some-utf16-file

The -a flag forces ripgrep to treat binary files as text, preventing it from skipping the file due to null bytes or other binary indicators.

Where Encoding Handling is Implemented

The encoding pipeline in ripgrep involves two critical source files:

  • src/core/main.rs: Contains the command-line parsing logic for -E/--encoding and orchestrates the transcoding step at runtime. This is where ripgrep decides whether to use BOM detection or user-specified encoding.

  • crates/ignore/src/encoding.rs: Houses the helper functions that map encoding names (like latin1 or gbk) to the underlying encoding_rs library implementations. This separation allows ripgrep to support dozens of legacy encodings through a unified interface.

Summary

  • Use -E/--encoding followed by the encoding name (e.g., latin1, gbk, shift_jis) to search files in that specific encoding.
  • ripgrep transcodes specified encodings to UTF-8 before regex matching, enabling full Unicode regex support on legacy files.
  • Specify -E none to disable all encoding logic and BOM sniffing, performing raw byte searches instead.
  • UTF-16 files with a BOM are automatically detected unless you use -E none.
  • Implementation resides in src/core/main.rs and crates/ignore/src/encoding.rs, using the encoding_rs library for conversions.

Frequently Asked Questions

Can ripgrep automatically detect file encodings?

No, ripgrep does not perform automatic encoding detection beyond UTF-16 BOM sniffing. You must explicitly specify the encoding using -E/--encoding for non-UTF-8 files like Latin-1 or GBK. Without the flag, ripgrep assumes UTF-8 and may miss matches or report encoding errors.

What happens when I search a UTF-16 file without specifying an encoding?

If the UTF-16 file contains a BOM (Byte Order Mark), ripgrep automatically detects and transcodes it to UTF-8 before searching, regardless of the -E flag. However, UTF-16 files without a BOM will be treated as UTF-8 (likely producing errors or missed matches) unless you explicitly specify the encoding or use -E none for raw-byte inspection.

How do I search a directory containing mixed encodings?

ripgrep applies a single encoding to all files in a search when using -E/--encoding. For directories with mixed encodings (e.g., some UTF-8, some GBK), you must run separate searches for each encoding, or use -E none to search raw bytes if you know the byte patterns you seek. There is no built-in per-file encoding auto-detection.

Why do I need the -a flag when using -E none?

The -a (or --text) flag forces ripgrep to treat binary files as text. When you use -E none, ripgrep operates on raw bytes and may encounter null bytes or sequences that look like binary data, causing it to skip the file or treat it as binary. The -a flag ensures ripgrep searches these files as text streams.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →