# How LiteParse Handles PDF Fonts with Problematic Encoding: Detection in extract.rs

> Learn how LiteParse addresses problematic PDF font encoding in extract.rs. Discover its detection methods for TrueType and Type 1 fonts, ensuring accurate text extraction.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: internals
- Published: 2026-06-06

---

**LiteParse handles PDF fonts with problematic encoding by running character-by-character extraction in [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs) where the `is_buggy_font` helper flags known bad TrueType and Type 1 font patterns, and a second filter rejects glyphs in the control-character and Private-Use Area ranges.**

Extracting reliable text from PDFs often fails when document producers embed fonts that use problematic encoding. The `run-llama/liteparse` crate solves this directly in its Rust extraction engine inside [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). This article breaks down how LiteParse identifies these problematic fonts and prevents corrupt characters from reaching the final output stream.

## Where Buggy Font Detection Happens

LiteParse does not rely on document-wide heuristics. Instead, it evaluates fonts during the character-by-character extraction phase in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). This granular strategy isolates fallback behavior to specific font streams rather than penalizing the entire document.

## Identifying Buggy Fonts with `is_buggy_font`

The core gatekeeper is the `is_buggy_font` helper found at lines 501–515 of [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). It inspects each font’s **PostScript name** and its `FontType` enum to decide whether a given font is a known offender.

### TrueType Subset Patterns

**TrueType** subset fonts that carry unreliable encoding tables are flagged when their PostScript name starts with `TT` or contains the `+TT` substring. These markers correlate with PDF generators that remap glyph indices in ways that break standard Unicode extraction.

### Type 1 Prefix and Underscore Pattern

LiteParse also flags **Type 1** fonts as buggy when their PostScript name carries a six-character prefix immediately followed by an underscore. This rigid naming convention identifies producers known to embed non-standard encoding vectors.

## Filtering Invalid Code Point Ranges

Even after a font passes the name check, individual characters must clear a second validation stage. Some PDF producers pack glyph indices into Unicode ranges that do not represent printable text.

LiteParse treats two bands as symptoms of buggy encoding:

- **Control characters** in the range `0x00` through `0x1F`
- **Private-Use Area** code points from `U+E000` to `U+F8FF`

When a glyph maps to either band, the extraction pipeline treats it as a buggy encoding artifact rather than valid content. This logic is expressed alongside the font classification routines in [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs).

```rust
// Conceptual validation based on crates/liteparse/src/extract.rs
fn is_buggy_encoding(c: char) -> bool {
    let cp = c as u32;
    cp <= 0x1F || (0xE000..=0xF8FF).contains(&cp)
}

```

## Why the Two-Stage Approach Matters

By combining font-level detection in `is_buggy_font` with character-level code point filtering, LiteParse builds a precise defense against bad PDF producers. The pipeline can apply targeted mitigation only when a known buggy font emits a suspicious glyph. This leaves correctly encoded text untouched while isolating the damage caused by non-compliant font encodings.

## Summary

- LiteParse detects PDF fonts with problematic encoding character-by-character in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).
- The `is_buggy_font` helper at lines 501–515 matches TrueType subset prefixes (`TT`, `+TT`) and Type 1 six-character-prefix-plus-underscore names.
- Glyphs mapped to control characters (≤ `0x1F`) or the Private-Use Area (`U+E000`–`U+F8FF`) are rejected as buggy artifacts.
- This two-stage validation preserves clean text while isolating damage from non-compliant font encodings.

## Frequently Asked Questions

### What source file in LiteParse handles buggy PDF font detection?

The detection logic lives in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). The `is_buggy_font` function at lines 501–515 performs the font name and type checks, while adjacent code point validation filters out individual buggy glyphs.

### How does LiteParse recognize a buggy TrueType font?

LiteParse flags a TrueType subset font as buggy if its PostScript name starts with `TT` or contains the substring `+TT`. These patterns indicate subset fonts produced by generators with known encoding defects.

### Which Unicode ranges does LiteParse treat as buggy during extraction?

LiteParse considers characters in the control-character range (`0x00`–`0x1F`) and the Private-Use Area (`U+E000`–`U+F8FF`) to be symptoms of problematic encoding. Glyphs in these ranges are discarded or handled as artifacts rather than valid text.

### Why does LiteParse check fonts character-by-character instead of scanning the whole document?

The character-by-character approach in [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs) localizes corrective behavior to specific font streams. This ensures that glyphs produced by known buggy font patterns trigger mitigation without affecting compliant text elsewhere in the PDF.