# How CMap and ToUnicode Parsing Works for CID Fonts in pdf‑inspector

> Learn how pdf-inspector parses CMap and ToUnicode for CID fonts. It converts CID codes to Unicode strings via embedded streams, TrueType fallback mappings, and glyph ID remapping for accurate text extraction.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-08

---

**pdf‑inspector converts CID codes to Unicode strings using a three-stage pipeline that parses embedded ToUnicode streams, builds fallback mappings from TrueType data when necessary, and remaps subset font glyph IDs to ensure accurate text extraction.**

Extracting readable text from PDFs containing **CID-keyed fonts** requires converting internal character identifiers into standard Unicode. The `pdf‑inspector` Rust library handles this through a robust **CMap and ToUnicode parsing** system implemented in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs). This article examines the exact mechanisms used to decode CID fonts, from parsing binary CMap streams to handling malformed resources through intelligent fallback strategies.

## The Three-Stage Processing Pipeline

The conversion process operates through three distinct stages defined in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs), each handling specific edge cases in PDF font encoding.

### Stage 1: Loading and Parsing ToUnicode CMaps

When a font dictionary contains a `/ToUnicode` stream, the raw bytes are fed to `ToUnicodeCMap::parse` (lines 1‑25). This parser handles both **binary** and **text** CMap formats automatically, detecting the byte width of source codes (1 or 2 bytes) from the CMap’s `codespace` definition. If the CMap is sparse or malformed, the system logs a warning at line 48 and triggers fallback handling rather than failing extraction.

### Stage 2: Building Fallback Mappings

If the PDF lacks a ToUnicode entry or the CMap proves unsuitable, the extractor attempts multiple **TrueType fallback** strategies (lines 944‑972). First, `build_cmap_from_truetype` reads the embedded font’s `cmap` table using `ttf_parser` to extract CID-to-Unicode mappings. For legacy Type 1 fonts, `build_simple_cmap_from_truetype` creates a minimal 1-byte mapping based on glyph order. When glyph names are available in the `post` table, `build_cmap_from_glyph_names` synthesizes the mapping. Finally, `build_cmap_from_builtin_cmap` loads pre-compiled binary CMaps such as **Adobe‑Korea1** (lines 1174‑1184) shipped with the library.

### Stage 3: CID-to-GID Remapping for Subset Fonts

Subset fonts often embed a `/CIDToGIDMap` that remaps CIDs to new sequential glyph IDs starting at 1. When detected, the system calls `remap_to_sequential` (lines 541‑565) to rebuild the CMap with corrected source codes. This ensures that the final Unicode output matches the actual embedded glyph data rather than referencing obsolete GIDs from the original font file.

## Character Lookup and Resolution

Once the CMap is prepared, the extractor translates content stream bytes through the `lookup` method (lines 436‑450). This function performs a per-byte CMap lookup, returning the appropriate Unicode string for each character code. The implementation handles both single-byte and double-byte source codes based on the width detected during the initial parse phase.

## Integration with the Font Extraction Pipeline

CMap construction is orchestrated within [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) (lines 1956‑1970) when building `FontCMaps` for each page’s font resources. The detector module ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)) determines whether CMap parsing is required based on the PDF type, while [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) consumes the resulting Unicode strings to generate final Markdown output.

## Practical Code Examples

The following patterns demonstrate how to use the public API for custom extraction workflows:

```rust
use pdf_inspector::tounicode::ToUnicodeCMap;

// Parse an embedded ToUnicode stream from PDF font resources
let raw_cmap: &[u8] = /* bytes from /ToUnicode stream */;
let cmap = ToUnicodeCMap::parse(raw_cmap)
    .expect("Failed to parse ToUnicode CMap");

// Build fallback from embedded TrueType font data
let truetype_data: &[u8] = /* font file bytes */;
let fallback_cmap = ToUnicodeCMap::build_cmap_from_truetype(truetype_data)
    .unwrap_or_else(|| ToUnicodeCMap::new());

// Remap when CIDToGIDMap is present (subset fonts)
let remapped = cmap.remap_to_sequential();

```

## Summary

- **ToUnicode CMap parsing** automatically handles binary and text formats, detecting 1-byte or 2-byte codespace widths during initialization.
- **Four fallback strategies**—TrueType `cmap` tables, simple glyph order mappings, glyph-name synthesis, and built-in Adobe CMaps—ensure text extraction succeeds even when ToUnicode entries are missing.
- **CID-to-GID remapping** corrects glyph ID mismatches in subset fonts by rebuilding the CMap with sequential GIDs starting at 1.
- The `lookup` method provides the final translation from processed CIDs to Unicode strings consumed by the Markdown converter.

## Frequently Asked Questions

### What happens when a PDF font lacks a ToUnicode entry?

pdf‑inspector attempts multiple fallback strategies, including reading the TrueType `cmap` table, parsing glyph names from the `post` table, or loading built-in CMaps like Adobe‑Korea1. These methods reconstruct the CID-to-Unicode mapping without requiring explicit ToUnicode data in the PDF.

### How does pdf‑inspector handle binary versus text CMap formats?

The parser in [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) automatically detects the format during the `ToUnicodeCMap::parse` operation. It processes binary streams through `BinaryCMapStream` and text streams through `EncodingCMap`, extracting the codespace definition to determine whether source codes are 1-byte or 2-byte wide.

### Why is CID-to-GID remapping necessary for subset fonts?

Subset fonts often embed a `/CIDToGIDMap` that reindexes glyph IDs to sequential values starting at 1. The `remap_to_sequential` function adjusts the CMap to reference these new GIDs, ensuring that the Unicode output aligns with the actual embedded glyph data rather than the original font's glyph IDs.

### Where does CMap integration occur in the extraction pipeline?

Font CMaps are constructed in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) (lines 1956‑1970) when processing each font resource, then used by [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) to generate the final Markdown output from the decoded Unicode strings.