# How to Search for Non-UTF-8 Encoded Files with ripgrep: Latin-1, GBK, and More

> Learn to search non UTF-8 encoded files like Latin-1 and GBK with ripgrep. Master RFC 1345, GBK, and raw byte searches with the -E flag.

- Repository: [Andrew Gallant/ripgrep](https://github.com/BurntSushi/ripgrep)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Use the `-E/--encoding` flag to specify the source encoding (e.g., `rg -E latin1` or `rg -E gbk`), which transcodes files to UTF-8 before searching, or use `-E none` to search raw bytes.**

By default, the `ripgrep` command-line tool assumes every file is UTF-8 encoded. However, legacy systems and international projects often rely on alternative encodings like Latin-1 (ISO-8859-1), GBK, or Shift_JIS. According to the BurntSushi/ripgrep source code, you can search for non-UTF-8 encoded files by leveraging the built-in transcoding pipeline that converts source encodings to UTF-8 on the fly.

## Understanding ripgrep’s Default Encoding Behavior

ripgrep operates on UTF-8 by default, but it includes a **BOM (Byte Order Mark) sniffing** mechanism. When it encounters a UTF-16 file with a BOM, it automatically handles transcoding regardless of command-line settings. This behavior is documented in [`GUIDE.md`](https://github.com/BurntSushi/ripgrep/blob/main/GUIDE.md) and implemented in [`src/core/main.rs`](https://github.com/BurntSushi/ripgrep/blob/main/src/core/main.rs), where the search logic first checks for UTF-16 BOM markers before applying user-specified encoding rules.

## Using the `-E/--encoding` Flag to Search Non-UTF-8 Files

The `-E` or `--encoding` flag tells ripgrep to treat every file it searches as being in the specified encoding. According to [`crates/ignore/src/encoding.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/encoding.rs), ripgrep uses the `encoding_rs` library to map encoding names like `latin1`, `gbk`, `euc-jp`, or `shift_jis` to their respective character sets.

When you specify an encoding, ripgrep transcodes the file contents from that encoding to UTF-8 before feeding them to the regex engine. This allows Unicode-aware regex features—such as `\w` for word characters or `.` for any character—to work correctly on non-UTF-8 source files.

### Searching Latin-1 (ISO-8859-1) Files

For Western European text files using single-byte Latin-1 encoding, specify the encoding explicitly:

```bash
rg "café" -E latin1 src/

```

ripgrep converts the Latin-1 bytes to UTF-8 internally, ensuring accented characters match correctly against your UTF-8 pattern.

### Searching GBK and Other Multi-Byte Encodings

For Chinese text encoded in GBK, or Japanese text in Shift_JIS or EUC-JP, use the appropriate encoding label:

```bash

# Search for the string "错误" in GBK-encoded files

rg "错误" -E gbk path/to/file.txt

```

The transcoding layer handles the conversion from GBK to UTF-8, allowing the regex engine to process the content as valid Unicode.

### Raw-Byte Searches with `-E none`

To bypass all encoding detection and transcoding—including BOM sniffing—use the special value `none`. This mode searches the file's raw bytes without conversion, which is useful for inspecting exact byte sequences or handling malformed files.

```bash

# Find exact UTF-16 byte sequences by searching raw bytes

rg '(?-u)\(\x045\x04@\x04;\x04>\x04:\x04' -E none -a some-utf16-file

```

The `-a` flag forces ripgrep to treat binary files as text, preventing it from skipping the file due to null bytes or other binary indicators.

## Where Encoding Handling is Implemented

The encoding pipeline in ripgrep involves two critical source files:

- **[`src/core/main.rs`](https://github.com/BurntSushi/ripgrep/blob/main/src/core/main.rs)**: Contains the command-line parsing logic for `-E/--encoding` and orchestrates the transcoding step at runtime. This is where ripgrep decides whether to use BOM detection or user-specified encoding.

- **[`crates/ignore/src/encoding.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/encoding.rs)**: Houses the helper functions that map encoding names (like `latin1` or `gbk`) to the underlying `encoding_rs` library implementations. This separation allows ripgrep to support dozens of legacy encodings through a unified interface.

## Summary

- Use **`-E/--encoding`** followed by the encoding name (e.g., `latin1`, `gbk`, `shift_jis`) to search files in that specific encoding.
- ripgrep **transcodes** specified encodings to UTF-8 before regex matching, enabling full Unicode regex support on legacy files.
- Specify **`-E none`** to disable all encoding logic and BOM sniffing, performing raw byte searches instead.
- UTF-16 files with a BOM are automatically detected unless you use `-E none`.
- Implementation resides in [`src/core/main.rs`](https://github.com/BurntSushi/ripgrep/blob/main/src/core/main.rs) and [`crates/ignore/src/encoding.rs`](https://github.com/BurntSushi/ripgrep/blob/main/crates/ignore/src/encoding.rs), using the `encoding_rs` library for conversions.

## Frequently Asked Questions

### Can ripgrep automatically detect file encodings?

No, ripgrep does not perform automatic encoding detection beyond **UTF-16 BOM sniffing**. You must explicitly specify the encoding using `-E/--encoding` for non-UTF-8 files like Latin-1 or GBK. Without the flag, ripgrep assumes UTF-8 and may miss matches or report encoding errors.

### What happens when I search a UTF-16 file without specifying an encoding?

If the UTF-16 file contains a **BOM (Byte Order Mark)**, ripgrep automatically detects and transcodes it to UTF-8 before searching, regardless of the `-E` flag. However, UTF-16 files without a BOM will be treated as UTF-8 (likely producing errors or missed matches) unless you explicitly specify the encoding or use `-E none` for raw-byte inspection.

### How do I search a directory containing mixed encodings?

ripgrep applies a **single encoding** to all files in a search when using `-E/--encoding`. For directories with mixed encodings (e.g., some UTF-8, some GBK), you must run separate searches for each encoding, or use `-E none` to search raw bytes if you know the byte patterns you seek. There is no built-in per-file encoding auto-detection.

### Why do I need the `-a` flag when using `-E none`?

The `-a` (or `--text`) flag forces ripgrep to treat binary files as text. When you use `-E none`, ripgrep operates on raw bytes and may encounter null bytes or sequences that look like binary data, causing it to skip the file or treat it as binary. The `-a` flag ensures ripgrep searches these files as text streams.