How pdf-inspector Handles ToUnicode CMap Parsing for CID-Encoded Fonts Like Identity-H
pdf-inspector handles CID-encoded fonts such as Identity-H by parsing ToUnicode CMap streams into a lookup table, then applying multiple fallback strategies—including encoding-based reconstruction, TrueType cmap parsing, and direct CID-to-Unicode passthrough—to recover Unicode text.
pdf-inspector is an open-source Rust library for extracting Unicode text from PDF documents. Its ToUnicodeCMap parser in src/tounicode.rs maps Character IDs (CIDs) to Unicode strings, which is especially critical for CID-encoded fonts like Identity-H where character codes do not correspond to standard single-byte encodings.
Core ToUnicodeCMap Structure
In src/tounicode.rs, the library defines the ToUnicodeCMap struct to store parsed mappings and metadata:
pub struct ToUnicodeCMap {
pub char_map: HashMap<u16, String>, // direct CID → Unicode
pub ranges: Vec<(u16, u16, u32)>, // CID range → base Unicode
pub code_byte_length: u8, // 1‑ or 2‑byte source codes
pub cid_passthrough: bool, // treat CID as Unicode when needed
}
(source: tounicode.rs lines 19–30)
The char_map field holds direct CID-to-Unicode mappings extracted from bfchar entries, while ranges stores interval mappings from bfrange sections. The code_byte_length field tracks whether source codes are one or two bytes, and cid_passthrough enables a last-resort fallback for fonts like Identity-H that encode Unicode values directly as CIDs.
Parsing the ToUnicode CMap Stream
The ToUnicodeCMap::parse method in src/tounicode.rs reads a decompressed ToUnicode stream and populates the struct. According to the pdf-inspector source code, the parser performs three main tasks across lines 110–165, 172–207, and 241–280.
Detecting Codespace Ranges
First, the parser determines the code_byte_length by reading the begincodespacerange section. If the PDF omits this codespace definition, pdf-inspector falls back to heuristics to infer the correct byte length.
Extracting bfchar and bfrange Entries
Next, it scans for bfchar entries in <src> <dst> format and bfrange entries in <start> <end> <base> or array form. These populate char_map and ranges, respectively.
Handling usecmap Directives
Finally, the parser respects the optional usecmap directive by merging in built-in binary CMaps when the stream references an external CMap resource.
Fallback Strategies for Missing or Insufficient CMaps
When a PDF lacks a ToUnicode CMap—or when the existing one is incomplete—pdf-inspector tries several reconstruction strategies, all implemented in src/tounicode.rs.
Encoding-Based Fallback
The function build_fallback_tounicode_from_encoding constructs a CMap by mapping the font’s Encoding dictionary—such as Identity-H—to a built-in UCS-2 CMap. It then applies the font’s specific encoding parameters to generate valid CID-to-Unicode mappings.
TrueType Fallback
For embedded TrueType fonts, build_cmap_from_truetype parses the font’s internal cmap table to derive a CID-to-Unicode map. This is particularly important for Identity-H fonts because CID equals glyph ID (cid == gid), allowing the TrueType table to bridge the gap when no ToUnicode stream exists.
Simple-Font Fallback
The function build_simple_cmap_from_truetype handles single-byte fonts by using the TrueType cmap subtable—for example, MacRoman or Windows-Symbol—and performing glyph-name lookups when necessary.
(source: fallback logic spans lines 284–331, 336–374, and 380–452 in tounicode.rs)
CID-to-Unicode Passthrough for Identity-H
If no usable CMap or font table is available, pdf-inspector sets the cid_passthrough flag. During decoding, the decode_cids method treats the raw CID value as a Unicode code point. This passthrough matches PDFs that encoded Unicode directly into CID values, a common pattern with Identity-H fonts:
if self.cid_passthrough {
if let Some(ch) = char::from_u32(cid as u32) { … }
}
(source: lines 779–886 in tounicode.rs)
Handling Subset-Font Remapping
Some PDF generators subset-embed fonts and renumber glyph IDs (GIDs), which breaks the original ToUnicode CMap. The function try_remap_subset_cmap in src/tounicode.rs detects this mismatch by inspecting the font’s Encoding dictionary and the minimum source CID.
When a mismatch is found, pdf-inspector either:
- Applies a CID-to-GIDMap via
build_cmap_with_cid_to_gid_mapif the font dictionary includes one, or - Remaps the CMap to sequential CID numbers using
remap_to_sequential.
This remapping ensures that even poorly generated subset PDFs yield correct Unicode output.
(source: lines 562–623 in tounicode.rs)
Integration with the Text Extractor
The CMap logic plugs into the broader extraction pipeline in src/extractor/fonts.rs. The helper get_font_file2_obj_num (lines 863–904) locates the embedded TrueType font stream used for fallback parsing. During content stream processing, extract_text_from_operand (lines 1089–1250) calls decode_cids or the fallback paths to produce the final Unicode text.
Orchestration happens in src/lib.rs through the public process_pdf_with_options API, while src/detector.rs classifies the PDF and determines when CMap processing is required.
Practical Code Examples
You can interact with pdf-inspector’s ToUnicode CMap parsing programmatically or through its CLI.
Parsing a Raw ToUnicode Stream
use pdf_inspector::tounicode::ToUnicodeCMap;
// `stream_bytes` = decompressed content of the /ToUnicode object
let cmap = ToUnicodeCMap::parse(&stream_bytes)
.expect("failed to build CMap");
// Look up a CID (e.g. 57) → Unicode string
let unicode = cmap.lookup(57).unwrap_or_default();
println!("CID 57 → {}", unicode);
Decoding a Byte Slice from a PDF Content Stream
use pdf_inspector::tounicode::ToUnicodeCMap;
// Assume `cmap` was built as above
let raw_bytes = vec![0x00, 0x39, 0x00, 0x41]; // two 2‑byte CIDs
let text = cmap.decode_cids(&raw_bytes);
println!("Decoded text: {}", text);
Using the CLI
# Extract plain text (Unicode) from a PDF that uses Identity‑H
pdf2md --json my‑document.pdf
Summary
- pdf-inspector’s
ToUnicodeCMapstruct insrc/tounicode.rsstores direct mappings, ranges, byte lengths, and a passthrough flag for CID-encoded fonts. - The
parsemethod extractsbfchar,bfrange, andbegincodespacerangedata from raw ToUnicode streams, merging external CMaps viausecmap. - For Identity-H and similar CID fonts, the library applies encoding-based, TrueType
cmap, and simple-font fallbacks before resorting to CID-to-Unicode passthrough. - Subset-font remapping via
try_remap_subset_cmaprepairs broken GID mappings by applying a CID-to-GIDMap or sequential remapping. - The extractor in
src/extractor/fonts.rsorchestrates these routines throughget_font_file2_obj_numandextract_text_from_operand.
Frequently Asked Questions
How does pdf-inspector handle Identity-H fonts that lack a ToUnicode CMap?
When Identity-H fonts omit a ToUnicode CMap, pdf-inspector first tries to rebuild a mapping from the font’s Encoding dictionary or embedded TrueType cmap table. If those fail, it enables cid_passthrough, which treats each CID value as a direct Unicode code point during the decode_cids phase.
What is the role of the cid_passthrough flag in src/tounicode.rs?
The cid_passthrough flag signals that the parser should interpret raw CID values as Unicode code points. This last-resort mechanism is essential for PDFs that store Unicode directly inside CID-encoded streams—most commonly encountered with Identity-H fonts that provide no explicit mapping.
How does pdf-inspector fix broken Unicode in subset-embedded fonts?
The try_remap_subset_cmap function detects when a subset font has renumbered glyph IDs and broken the original CMap. It then either applies a CID-to-GIDMap from the font dictionary or remaps the CMap to sequential CID numbers via remap_to_sequential, restoring correct Unicode output.
Where does the text extractor invoke the ToUnicode CMap decoder?
In src/extractor/fonts.rs, the function extract_text_from_operand (lines 1089–1250) drives text extraction by calling decode_cids on the parsed ToUnicodeCMap. The helper get_font_file2_obj_num (lines 863–904) locates the embedded font stream needed for TrueType fallback parsing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →