# How pdf-inspector Handles ToUnicode CMap Parsing for CID-Encoded Fonts Like Identity-H

> Discover how pdf-inspector parses ToUnicode CMaps for CID-encoded fonts like Identity-H. Learn about its fallback strategies for accurate Unicode text recovery.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: internals
- Published: 2026-08-06

---

**pdf-inspector handles CID-encoded fonts such as Identity-H by parsing ToUnicode CMap streams into a lookup table, then applying multiple fallback strategies—including encoding-based reconstruction, TrueType cmap parsing, and direct CID-to-Unicode passthrough—to recover Unicode text.**

pdf-inspector is an open-source Rust library for extracting Unicode text from PDF documents. Its `ToUnicodeCMap` parser in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) maps Character IDs (CIDs) to Unicode strings, which is especially critical for CID-encoded fonts like *Identity-H* where character codes do not correspond to standard single-byte encodings.

## Core ToUnicodeCMap Structure

In [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), the library defines the **`ToUnicodeCMap`** struct to store parsed mappings and metadata:

```rust
pub struct ToUnicodeCMap {
    pub char_map: HashMap<u16, String>,   // direct CID → Unicode
    pub ranges:   Vec<(u16, u16, u32)>,   // CID range → base Unicode
    pub code_byte_length: u8,             // 1‑ or 2‑byte source codes
    pub cid_passthrough: bool,            // treat CID as Unicode when needed
}

```

*(source: [tounicode.rs lines 19–30](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs#L19-L30))*

The **`char_map`** field holds direct CID-to-Unicode mappings extracted from `bfchar` entries, while **`ranges`** stores interval mappings from `bfrange` sections. The **`code_byte_length`** field tracks whether source codes are one or two bytes, and **`cid_passthrough`** enables a last-resort fallback for fonts like *Identity-H* that encode Unicode values directly as CIDs.

## Parsing the ToUnicode CMap Stream

The **`ToUnicodeCMap::parse`** method in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) reads a decompressed ToUnicode stream and populates the struct. According to the pdf-inspector source code, the parser performs three main tasks across lines 110–165, 172–207, and 241–280.

### Detecting Codespace Ranges

First, the parser determines the **`code_byte_length`** by reading the `begincodespacerange` section. If the PDF omits this codespace definition, pdf-inspector falls back to heuristics to infer the correct byte length.

### Extracting bfchar and bfrange Entries

Next, it scans for **`bfchar`** entries in `<src> <dst>` format and **`bfrange`** entries in `<start> <end> <base>` or array form. These populate `char_map` and `ranges`, respectively.

### Handling usecmap Directives

Finally, the parser respects the optional **`usecmap`** directive by merging in built-in binary CMaps when the stream references an external CMap resource.

## Fallback Strategies for Missing or Insufficient CMaps

When a PDF lacks a ToUnicode CMap—or when the existing one is incomplete—pdf-inspector tries several reconstruction strategies, all implemented in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs).

### Encoding-Based Fallback

The function **`build_fallback_tounicode_from_encoding`** constructs a CMap by mapping the font’s *Encoding* dictionary—such as *Identity-H*—to a built-in UCS-2 CMap. It then applies the font’s specific encoding parameters to generate valid CID-to-Unicode mappings.

### TrueType Fallback

For embedded TrueType fonts, **`build_cmap_from_truetype`** parses the font’s internal `cmap` table to derive a CID-to-Unicode map. This is particularly important for *Identity-H* fonts because CID equals glyph ID (`cid == gid`), allowing the TrueType table to bridge the gap when no ToUnicode stream exists.

### Simple-Font Fallback

The function **`build_simple_cmap_from_truetype`** handles single-byte fonts by using the TrueType `cmap` subtable—for example, MacRoman or Windows-Symbol—and performing glyph-name lookups when necessary.

*(source: fallback logic spans lines 284–331, 336–374, and 380–452 in [tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs))*

### CID-to-Unicode Passthrough for Identity-H

If no usable CMap or font table is available, pdf-inspector sets the **`cid_passthrough`** flag. During decoding, the **`decode_cids`** method treats the raw CID value as a Unicode code point. This passthrough matches PDFs that encoded Unicode directly into CID values, a common pattern with *Identity-H* fonts:

```rust
if self.cid_passthrough {
    if let Some(ch) = char::from_u32(cid as u32) { … }
}

```

*(source: lines 779–886 in [tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs))*

## Handling Subset-Font Remapping

Some PDF generators subset-embed fonts and renumber glyph IDs (GIDs), which breaks the original ToUnicode CMap. The function **`try_remap_subset_cmap`** in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) detects this mismatch by inspecting the font’s *Encoding* dictionary and the minimum source CID.

When a mismatch is found, pdf-inspector either:

- Applies a **CID-to-GIDMap** via **`build_cmap_with_cid_to_gid_map`** if the font dictionary includes one, or
- **Remaps** the CMap to sequential CID numbers using **`remap_to_sequential`**.

This remapping ensures that even poorly generated subset PDFs yield correct Unicode output.

*(source: lines 562–623 in [tounicode.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs))*

## Integration with the Text Extractor

The CMap logic plugs into the broader extraction pipeline in **[`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)**. The helper **`get_font_file2_obj_num`** (lines 863–904) locates the embedded TrueType font stream used for fallback parsing. During content stream processing, **`extract_text_from_operand`** (lines 1089–1250) calls `decode_cids` or the fallback paths to produce the final Unicode text.

Orchestration happens in **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** through the public **`process_pdf_with_options`** API, while **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)** classifies the PDF and determines when CMap processing is required.

## Practical Code Examples

You can interact with pdf-inspector’s ToUnicode CMap parsing programmatically or through its CLI.

### Parsing a Raw ToUnicode Stream

```rust
use pdf_inspector::tounicode::ToUnicodeCMap;

// `stream_bytes` = decompressed content of the /ToUnicode object
let cmap = ToUnicodeCMap::parse(&stream_bytes)
    .expect("failed to build CMap");

// Look up a CID (e.g. 57) → Unicode string
let unicode = cmap.lookup(57).unwrap_or_default();
println!("CID 57 → {}", unicode);

```

### Decoding a Byte Slice from a PDF Content Stream

```rust
use pdf_inspector::tounicode::ToUnicodeCMap;

// Assume `cmap` was built as above
let raw_bytes = vec![0x00, 0x39, 0x00, 0x41]; // two 2‑byte CIDs
let text = cmap.decode_cids(&raw_bytes);
println!("Decoded text: {}", text);

```

### Using the CLI

```bash

# Extract plain text (Unicode) from a PDF that uses Identity‑H

pdf2md --json my‑document.pdf

```

## Summary

- pdf-inspector’s **`ToUnicodeCMap`** struct in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) stores direct mappings, ranges, byte lengths, and a passthrough flag for CID-encoded fonts.
- The **`parse`** method extracts `bfchar`, `bfrange`, and `begincodespacerange` data from raw ToUnicode streams, merging external CMaps via `usecmap`.
- For *Identity-H* and similar CID fonts, the library applies **encoding-based**, **TrueType `cmap`**, and **simple-font** fallbacks before resorting to **CID-to-Unicode passthrough**.
- **Subset-font remapping** via `try_remap_subset_cmap` repairs broken GID mappings by applying a CID-to-GIDMap or sequential remapping.
- The extractor in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) orchestrates these routines through **`get_font_file2_obj_num`** and **`extract_text_from_operand`**.

## Frequently Asked Questions

### How does pdf-inspector handle Identity-H fonts that lack a ToUnicode CMap?

When *Identity-H* fonts omit a ToUnicode CMap, pdf-inspector first tries to rebuild a mapping from the font’s *Encoding* dictionary or embedded TrueType `cmap` table. If those fail, it enables **`cid_passthrough`**, which treats each CID value as a direct Unicode code point during the `decode_cids` phase.

### What is the role of the `cid_passthrough` flag in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)?

The **`cid_passthrough`** flag signals that the parser should interpret raw CID values as Unicode code points. This last-resort mechanism is essential for PDFs that store Unicode directly inside CID-encoded streams—most commonly encountered with *Identity-H* fonts that provide no explicit mapping.

### How does pdf-inspector fix broken Unicode in subset-embedded fonts?

The **`try_remap_subset_cmap`** function detects when a subset font has renumbered glyph IDs and broken the original CMap. It then either applies a **CID-to-GIDMap** from the font dictionary or remaps the CMap to sequential CID numbers via **`remap_to_sequential`**, restoring correct Unicode output.

### Where does the text extractor invoke the ToUnicode CMap decoder?

In **[`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)**, the function **`extract_text_from_operand`** (lines 1089–1250) drives text extraction by calling `decode_cids` on the parsed `ToUnicodeCMap`. The helper **`get_font_file2_obj_num`** (lines 863–904) locates the embedded font stream needed for TrueType fallback parsing.