How pdf-inspector Handles ToUnicode CMap Parsing for CID-Encoded Fonts Like Identity-H

pdf-inspector handles CID-encoded fonts such as Identity-H by parsing ToUnicode CMap streams into a lookup table, then applying multiple fallback strategies—including encoding-based reconstruction, TrueType cmap parsing, and direct CID-to-Unicode passthrough—to recover Unicode text.

pdf-inspector is an open-source Rust library for extracting Unicode text from PDF documents. Its ToUnicodeCMap parser in src/tounicode.rs maps Character IDs (CIDs) to Unicode strings, which is especially critical for CID-encoded fonts like Identity-H where character codes do not correspond to standard single-byte encodings.

Core ToUnicodeCMap Structure

In src/tounicode.rs, the library defines the ToUnicodeCMap struct to store parsed mappings and metadata:

pub struct ToUnicodeCMap {
    pub char_map: HashMap<u16, String>,   // direct CID → Unicode
    pub ranges:   Vec<(u16, u16, u32)>,   // CID range → base Unicode
    pub code_byte_length: u8,             // 1‑ or 2‑byte source codes
    pub cid_passthrough: bool,            // treat CID as Unicode when needed
}

(source: tounicode.rs lines 19–30)

The char_map field holds direct CID-to-Unicode mappings extracted from bfchar entries, while ranges stores interval mappings from bfrange sections. The code_byte_length field tracks whether source codes are one or two bytes, and cid_passthrough enables a last-resort fallback for fonts like Identity-H that encode Unicode values directly as CIDs.

Parsing the ToUnicode CMap Stream

The ToUnicodeCMap::parse method in src/tounicode.rs reads a decompressed ToUnicode stream and populates the struct. According to the pdf-inspector source code, the parser performs three main tasks across lines 110–165, 172–207, and 241–280.

Detecting Codespace Ranges

First, the parser determines the code_byte_length by reading the begincodespacerange section. If the PDF omits this codespace definition, pdf-inspector falls back to heuristics to infer the correct byte length.

Extracting bfchar and bfrange Entries

Next, it scans for bfchar entries in <src> <dst> format and bfrange entries in <start> <end> <base> or array form. These populate char_map and ranges, respectively.

Handling usecmap Directives

Finally, the parser respects the optional usecmap directive by merging in built-in binary CMaps when the stream references an external CMap resource.

Fallback Strategies for Missing or Insufficient CMaps

When a PDF lacks a ToUnicode CMap—or when the existing one is incomplete—pdf-inspector tries several reconstruction strategies, all implemented in src/tounicode.rs.

Encoding-Based Fallback

The function build_fallback_tounicode_from_encoding constructs a CMap by mapping the font’s Encoding dictionary—such as Identity-H—to a built-in UCS-2 CMap. It then applies the font’s specific encoding parameters to generate valid CID-to-Unicode mappings.

TrueType Fallback

For embedded TrueType fonts, build_cmap_from_truetype parses the font’s internal cmap table to derive a CID-to-Unicode map. This is particularly important for Identity-H fonts because CID equals glyph ID (cid == gid), allowing the TrueType table to bridge the gap when no ToUnicode stream exists.

Simple-Font Fallback

The function build_simple_cmap_from_truetype handles single-byte fonts by using the TrueType cmap subtable—for example, MacRoman or Windows-Symbol—and performing glyph-name lookups when necessary.

(source: fallback logic spans lines 284–331, 336–374, and 380–452 in tounicode.rs)

CID-to-Unicode Passthrough for Identity-H

If no usable CMap or font table is available, pdf-inspector sets the cid_passthrough flag. During decoding, the decode_cids method treats the raw CID value as a Unicode code point. This passthrough matches PDFs that encoded Unicode directly into CID values, a common pattern with Identity-H fonts:

if self.cid_passthrough {
    if let Some(ch) = char::from_u32(cid as u32) { … }
}

(source: lines 779–886 in tounicode.rs)

Handling Subset-Font Remapping

Some PDF generators subset-embed fonts and renumber glyph IDs (GIDs), which breaks the original ToUnicode CMap. The function try_remap_subset_cmap in src/tounicode.rs detects this mismatch by inspecting the font’s Encoding dictionary and the minimum source CID.

When a mismatch is found, pdf-inspector either:

  • Applies a CID-to-GIDMap via build_cmap_with_cid_to_gid_map if the font dictionary includes one, or
  • Remaps the CMap to sequential CID numbers using remap_to_sequential.

This remapping ensures that even poorly generated subset PDFs yield correct Unicode output.

(source: lines 562–623 in tounicode.rs)

Integration with the Text Extractor

The CMap logic plugs into the broader extraction pipeline in src/extractor/fonts.rs. The helper get_font_file2_obj_num (lines 863–904) locates the embedded TrueType font stream used for fallback parsing. During content stream processing, extract_text_from_operand (lines 1089–1250) calls decode_cids or the fallback paths to produce the final Unicode text.

Orchestration happens in src/lib.rs through the public process_pdf_with_options API, while src/detector.rs classifies the PDF and determines when CMap processing is required.

Practical Code Examples

You can interact with pdf-inspector’s ToUnicode CMap parsing programmatically or through its CLI.

Parsing a Raw ToUnicode Stream

use pdf_inspector::tounicode::ToUnicodeCMap;

// `stream_bytes` = decompressed content of the /ToUnicode object
let cmap = ToUnicodeCMap::parse(&stream_bytes)
    .expect("failed to build CMap");

// Look up a CID (e.g. 57) → Unicode string
let unicode = cmap.lookup(57).unwrap_or_default();
println!("CID 57 → {}", unicode);

Decoding a Byte Slice from a PDF Content Stream

use pdf_inspector::tounicode::ToUnicodeCMap;

// Assume `cmap` was built as above
let raw_bytes = vec![0x00, 0x39, 0x00, 0x41]; // two 2‑byte CIDs
let text = cmap.decode_cids(&raw_bytes);
println!("Decoded text: {}", text);

Using the CLI


# Extract plain text (Unicode) from a PDF that uses Identity‑H

pdf2md --json my‑document.pdf

Summary

  • pdf-inspector’s ToUnicodeCMap struct in src/tounicode.rs stores direct mappings, ranges, byte lengths, and a passthrough flag for CID-encoded fonts.
  • The parse method extracts bfchar, bfrange, and begincodespacerange data from raw ToUnicode streams, merging external CMaps via usecmap.
  • For Identity-H and similar CID fonts, the library applies encoding-based, TrueType cmap, and simple-font fallbacks before resorting to CID-to-Unicode passthrough.
  • Subset-font remapping via try_remap_subset_cmap repairs broken GID mappings by applying a CID-to-GIDMap or sequential remapping.
  • The extractor in src/extractor/fonts.rs orchestrates these routines through get_font_file2_obj_num and extract_text_from_operand.

Frequently Asked Questions

How does pdf-inspector handle Identity-H fonts that lack a ToUnicode CMap?

When Identity-H fonts omit a ToUnicode CMap, pdf-inspector first tries to rebuild a mapping from the font’s Encoding dictionary or embedded TrueType cmap table. If those fail, it enables cid_passthrough, which treats each CID value as a direct Unicode code point during the decode_cids phase.

What is the role of the cid_passthrough flag in src/tounicode.rs?

The cid_passthrough flag signals that the parser should interpret raw CID values as Unicode code points. This last-resort mechanism is essential for PDFs that store Unicode directly inside CID-encoded streams—most commonly encountered with Identity-H fonts that provide no explicit mapping.

How does pdf-inspector fix broken Unicode in subset-embedded fonts?

The try_remap_subset_cmap function detects when a subset font has renumbered glyph IDs and broken the original CMap. It then either applies a CID-to-GIDMap from the font dictionary or remaps the CMap to sequential CID numbers via remap_to_sequential, restoring correct Unicode output.

Where does the text extractor invoke the ToUnicode CMap decoder?

In src/extractor/fonts.rs, the function extract_text_from_operand (lines 1089–1250) drives text extraction by calling decode_cids on the parsed ToUnicodeCMap. The helper get_font_file2_obj_num (lines 863–904) locates the embedded font stream needed for TrueType fallback parsing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →