How CMap and ToUnicode Parsing Works for CID Fonts in pdf‑inspector
pdf‑inspector converts CID codes to Unicode strings using a three-stage pipeline that parses embedded ToUnicode streams, builds fallback mappings from TrueType data when necessary, and remaps subset font glyph IDs to ensure accurate text extraction.
Extracting readable text from PDFs containing CID-keyed fonts requires converting internal character identifiers into standard Unicode. The pdf‑inspector Rust library handles this through a robust CMap and ToUnicode parsing system implemented in src/tounicode.rs. This article examines the exact mechanisms used to decode CID fonts, from parsing binary CMap streams to handling malformed resources through intelligent fallback strategies.
The Three-Stage Processing Pipeline
The conversion process operates through three distinct stages defined in src/tounicode.rs, each handling specific edge cases in PDF font encoding.
Stage 1: Loading and Parsing ToUnicode CMaps
When a font dictionary contains a /ToUnicode stream, the raw bytes are fed to ToUnicodeCMap::parse (lines 1‑25). This parser handles both binary and text CMap formats automatically, detecting the byte width of source codes (1 or 2 bytes) from the CMap’s codespace definition. If the CMap is sparse or malformed, the system logs a warning at line 48 and triggers fallback handling rather than failing extraction.
Stage 2: Building Fallback Mappings
If the PDF lacks a ToUnicode entry or the CMap proves unsuitable, the extractor attempts multiple TrueType fallback strategies (lines 944‑972). First, build_cmap_from_truetype reads the embedded font’s cmap table using ttf_parser to extract CID-to-Unicode mappings. For legacy Type 1 fonts, build_simple_cmap_from_truetype creates a minimal 1-byte mapping based on glyph order. When glyph names are available in the post table, build_cmap_from_glyph_names synthesizes the mapping. Finally, build_cmap_from_builtin_cmap loads pre-compiled binary CMaps such as Adobe‑Korea1 (lines 1174‑1184) shipped with the library.
Stage 3: CID-to-GID Remapping for Subset Fonts
Subset fonts often embed a /CIDToGIDMap that remaps CIDs to new sequential glyph IDs starting at 1. When detected, the system calls remap_to_sequential (lines 541‑565) to rebuild the CMap with corrected source codes. This ensures that the final Unicode output matches the actual embedded glyph data rather than referencing obsolete GIDs from the original font file.
Character Lookup and Resolution
Once the CMap is prepared, the extractor translates content stream bytes through the lookup method (lines 436‑450). This function performs a per-byte CMap lookup, returning the appropriate Unicode string for each character code. The implementation handles both single-byte and double-byte source codes based on the width detected during the initial parse phase.
Integration with the Font Extraction Pipeline
CMap construction is orchestrated within src/extractor/fonts.rs (lines 1956‑1970) when building FontCMaps for each page’s font resources. The detector module (src/detector.rs) determines whether CMap parsing is required based on the PDF type, while src/markdown/convert.rs consumes the resulting Unicode strings to generate final Markdown output.
Practical Code Examples
The following patterns demonstrate how to use the public API for custom extraction workflows:
use pdf_inspector::tounicode::ToUnicodeCMap;
// Parse an embedded ToUnicode stream from PDF font resources
let raw_cmap: &[u8] = /* bytes from /ToUnicode stream */;
let cmap = ToUnicodeCMap::parse(raw_cmap)
.expect("Failed to parse ToUnicode CMap");
// Build fallback from embedded TrueType font data
let truetype_data: &[u8] = /* font file bytes */;
let fallback_cmap = ToUnicodeCMap::build_cmap_from_truetype(truetype_data)
.unwrap_or_else(|| ToUnicodeCMap::new());
// Remap when CIDToGIDMap is present (subset fonts)
let remapped = cmap.remap_to_sequential();
Summary
- ToUnicode CMap parsing automatically handles binary and text formats, detecting 1-byte or 2-byte codespace widths during initialization.
- Four fallback strategies—TrueType
cmaptables, simple glyph order mappings, glyph-name synthesis, and built-in Adobe CMaps—ensure text extraction succeeds even when ToUnicode entries are missing. - CID-to-GID remapping corrects glyph ID mismatches in subset fonts by rebuilding the CMap with sequential GIDs starting at 1.
- The
lookupmethod provides the final translation from processed CIDs to Unicode strings consumed by the Markdown converter.
Frequently Asked Questions
What happens when a PDF font lacks a ToUnicode entry?
pdf‑inspector attempts multiple fallback strategies, including reading the TrueType cmap table, parsing glyph names from the post table, or loading built-in CMaps like Adobe‑Korea1. These methods reconstruct the CID-to-Unicode mapping without requiring explicit ToUnicode data in the PDF.
How does pdf‑inspector handle binary versus text CMap formats?
The parser in src/tounicode.rs automatically detects the format during the ToUnicodeCMap::parse operation. It processes binary streams through BinaryCMapStream and text streams through EncodingCMap, extracting the codespace definition to determine whether source codes are 1-byte or 2-byte wide.
Why is CID-to-GID remapping necessary for subset fonts?
Subset fonts often embed a /CIDToGIDMap that reindexes glyph IDs to sequential values starting at 1. The remap_to_sequential function adjusts the CMap to reference these new GIDs, ensuring that the Unicode output aligns with the actual embedded glyph data rather than the original font's glyph IDs.
Where does CMap integration occur in the extraction pipeline?
Font CMaps are constructed in src/extractor/fonts.rs (lines 1956‑1970) when processing each font resource, then used by src/markdown/convert.rs to generate the final Markdown output from the decoded Unicode strings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →