# Handling CID Fonts with ToUnicode CMap Decoding in pdf-inspector: A Complete Guide

> Learn to extract Unicode text from CID fonts in pdf-inspector using ToUnicode CMap decoding. Our guide covers fallback mechanisms for missing mapping data.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-10

---

**`pdf-inspector` extracts Unicode text from CID-encoded PDF fonts by parsing `/ToUnicode` streams with automatic fallback to TrueType cmap tables, glyph names, and built-in CMaps when mapping data is missing or incomplete.**

The `firecrawl/pdf-inspector` repository implements a robust multi-layer strategy for decoding **CID fonts** (Character Identifier fonts) commonly found in PDF documents. CID fonts use numeric identifiers rather than direct Unicode codepoints, making **ToUnicode CMap decoding** essential for accurate text extraction. This article explains the internal pipeline, fallback mechanisms, and practical implementation details drawn directly from the source code.

## What Are CID Fonts and ToUnicode CMaps?

**CID fonts** are font formats where each glyph is referenced by a **Character Identifier (CID)**—a 16-bit number—rather than a character name or Unicode value. When you open a PDF containing Chinese, Japanese, or Korean text, or certain composite Latin fonts, you're almost certainly dealing with CID fonts.

The **ToUnicode CMap** is a PDF object that maps these CIDs to actual Unicode strings. Without it, raw CID values are meaningless to text extraction tools. According to the `pdf-inspector` source, the core decoding logic resides in **[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)**, orchestrated by the font extraction layer in **[`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)**.

## The Six-Stage ToUnicode Pipeline

The `pdf-inspector` crate processes CID font decoding through a carefully ordered pipeline. Each stage attempts to build a valid CID-to-Unicode mapping, with subsequent stages serving as fallbacks.

| Stage | Function | Purpose | Trigger Condition |
|-------|----------|---------|-------------------|
| 1 | `ToUnicodeCMap::parse` | Parse explicit `/ToUnicode` stream | `/ToUnicode` key present in font dictionary |
| 2 | `build_cmap_from_truetype` | Extract cmap table from embedded TrueType font | ToUnicode sparse or missing; TrueType data available |
| 3 | `build_cmap_from_glyph_names` | Convert glyph names via Adobe Glyph List | TrueType cmap unavailable; *post* table present |
| 4 | `build_cmap_from_builtin_cmap` | Load legacy CID→Unicode binary tables | CIDSystemInfo matches known ordering (e.g., `Adobe-Identity`) |
| 5 | `merge_cmaps` | Combine primary map with richer fallback | Overlay contains mappings missing from base |
| 6 | `remap_to_sequential` | Reorder CIDs to sequential space | Simplifies downstream layout processing |

### Stage 1: Parsing Raw ToUnicode Streams

The entry point for explicit ToUnicode data is `ToUnicodeCMap::parse` at lines 103-110 of [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs):

```rust
pub fn parse(data: &[u8]) -> Result<ToUnicodeCMap, String> {
    let mut cmap = ToUnicodeCMap::new();
    // … parsing logic …
    Ok(cmap)
}

```

This function decompresses the stream, handles CMap syntax (begincidchar, beginbfchar, etc.), and populates a `ToUnicodeCMap` struct containing the CID→Unicode mappings.

### Stage 2: TrueType cmap Fallback

When the ToUnicode stream is missing or incomplete, `pdf-inspector` examines the embedded font program. The `build_cmap_from_truetype` function (lines 944-952) uses the `ttf_parser` crate to read the TrueType `cmap` table:

```rust
pub fn build_cmap_from_truetype(font_data: &[u8]) -> Option<ToUnicodeCMap> {
    let mut cmap = ToUnicodeCMap::new();
    // read the TrueType `cmap` table via `ttf_parser`
    // … populate `cmap` …
    Some(cmap)
}

```

This often succeeds where PDF ToUnicode fails, since font embedding is mandatory for subset CID fonts.

### Stage 3: Glyph Name Resolution

For fonts with neither ToUnicode nor usable TrueType cmap data, the extractor falls back to **glyph names** from the TrueType *post* table, converting them via the Adobe Glyph List:

```rust
fn build_cmap_from_glyph_names(face: &ttf_parser::Face<'_>) -> Option<ToUnicodeCMap> {
    let mut cmap = ToUnicodeCMap::new();
    // map glyph names → Unicode using `adobe_glyphlist`
    // …
    Some(cmap)
}

```

This handles older PDFs and certain technical drawing fonts where glyph naming conventions follow Adobe standards.

### Stage 4: Built-in Binary CMaps

Legacy CID fonts reference **CIDSystemInfo** dictionaries with `/Ordering` values like `Adobe-Japan1` or `Adobe-Identity`. The `build_cmap_from_builtin_cmap` function loads pre-compiled binary CMaps shipped with the crate for these known orderings.

### Stage 5 & 6: Merging and Remapping

The `merge_cmaps` function combines a sparse primary CMap with a richer fallback:

```rust
fn merge_cmaps(base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap

```

`remap_to_sequential` optionally reorders the CID space to eliminate gaps, simplifying the text layout engine's job.

## How the Font Extractor Orchestrates CMap Resolution

The **[`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs)** module coordinates all CMap building. At lines 340-380, the `resolve_font_cmap` function traverses each font dictionary:

```rust
pub fn resolve_font_cmap(font_dict: &Dict, resources: &Resources) -> CharMap {
    if let Some(cmap) = font_dict.get(b"ToUnicode") {
        // parse the stream
        let parsed = ToUnicodeCMap::parse(&cmap_data)?;
        // possibly merge with a fallback
        let final_map = merge_cmaps(parsed, fallback);
        char_map.insert(font_id, final_map);
    } else {
        // use TrueType or builtin fallback
        let fallback = build_cmap_from_truetype(...).or_else(...);
    }
    // …
}

```

The returned `CharMap` is consumed by the text extraction engine to convert raw character codes to Unicode strings during page content parsing.

### OCR Flagging for Unrecoverable Fonts

When **no mapping can be constructed**—common with *Identity-H* fonts lacking any usable cmap—the extractor marks the page as needing OCR. This logic appears around line 1594 in [`tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tounicode.rs), with downstream consumers in [`detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detector.rs) surfacing the flag to users.

## Handling Edge Cases and Broken PDFs

The `pdf-inspector` codebase contains specific mitigations for real-world PDF problems:

- **Sparse ToUnicode streams**: Logged warning at line 48; TrueType cmap preferred when richer
- **Broken subset fonts**: Detection and repair logic at line 851 ("detect and fix broken ToUnicode CMaps")
- **Identity-H/V without mapping**: Forced fallback through the three-stage cascade rather than emitting empty text

## Practical Usage: Extracting Text from CID Font PDFs

Below is a complete Rust example using `pdf-inspector`'s public API. The ToUnicode decoding happens automatically during `process_pdf`:

```rust
use pdf_inspector::{process_pdf, MarkdownProfile};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Path to a PDF that uses CID (Identity‑H) fonts
    let pdf_path = "samples/cid-fonts.pdf";

    // Process the PDF with the default profile (fidelity)
    let result = process_pdf(pdf_path)?;

    // `result` contains a vector of `TextLine` items already decoded via ToUnicode CMaps
    for line in result.items.iter() {
        println!("{}", line.text);
    }

    Ok(())
}

```

The `process_pdf` function delegates to the extractor, which runs the full CMap resolution pipeline described above. Unicode text extraction succeeds even when the original PDF lacks explicit ToUnicode data.

## Key Source Files Reference

| File | Role | Key Symbols |
|------|------|-------------|
| [`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs) | Core parser, fallback implementations, merging | `ToUnicodeCMap::parse`, `build_cmap_from_truetype`, `merge_cmaps` |
| [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | Font dictionary traversal, CMap selection | `resolve_font_cmap`, `CharMap` |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API entry points | `process_pdf`, `extract_text` |
| [`tests/integration_tests.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tests/integration_tests.rs) | CID font validation, OCR flag testing | Integration test cases |

## Summary

- **[`src/tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tounicode.rs)** implements the complete ToUnicode CMap decoding pipeline with four fallback strategies for robust CID font handling
- **`ToUnicodeCMap::parse`** processes explicit `/ToUnicode` streams; **`build_cmap_from_truetype`** extracts mappings from embedded font programs
- **`resolve_font_cmap`** in [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) orchestrates fallback selection and exposes `CharMap` to the extraction engine
- Unrecoverable mappings trigger OCR flags rather than silent failure, ensuring users know when automated extraction failed
- The public `process_pdf` API automatically runs all decoding logic—no manual CMap handling required for typical use

## Frequently Asked Questions

### What happens if a PDF has CID fonts but no ToUnicode entry?

The extractor attempts three fallback strategies in order: **TrueType cmap table extraction**, **glyph name resolution via Adobe Glyph List**, and **built-in binary CMaps** for known CIDSystemInfo orderings. If all fail, the font is flagged for OCR rather than returning empty or garbage text.

### Why does Identity-H sometimes fail to extract text while other CID fonts work?

**Identity-H** fonts map CIDs directly to glyph IDs without a predefined character collection. Without an explicit ToUnicode CMap or embedded TrueType cmap, there's no semantic mapping available. The `pdf-inspector` detector specifically identifies this case (around line 1594 in [`tounicode.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tounicode.rs)) and triggers OCR fallback.

### Can I manually inspect the CMap selected for a specific font?

Currently, the `CharMap` resolution happens internally within `resolve_font_cmap`. For debugging, enable logging at the `DEBUG` level—`pdf-inspector` logs CMap source selection and fallback activation. The `ToUnicodeCMap` struct is not yet exposed in the public API, though internal unit tests in [`tests/integration_tests.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/tests/integration_tests.rs) demonstrate direct inspection patterns.

### How does the crate handle corrupted or partial ToUnicode streams?

The parser detects sparse mappings and logs warnings (line 48). It then prefers **TrueType cmap data** when available, as embedded fonts typically contain more complete mappings than PDF ToUnicode streams. For broken subset fonts with mismatched GIDs, reconstruction logic at line 851 attempts to rebuild a valid mapping from available glyph data.