How LiteParse Handles Font Metadata and Buggy Font Encoding Detection in PDFs
LiteParse extracts rich font metadata from PDF text objects using a Rust wrapper around PDFium and flags problematic encodings via the font_is_buggy boolean in the TextItem struct, enabling downstream consumers to handle malformed fonts gracefully.
LiteParse, the Rust-based PDF parsing library from the run-llama/liteparse repository, extracts typographic details and detects corrupted font encodings through a two-phase heuristic system. The library captures font names, types, and embedding status while analyzing individual glyphs for private-use codepoints and control characters. This metadata is exposed through the TextItem struct, allowing applications to identify text segments that may require special handling such as OCR fallback or custom glyph remapping.
Font Metadata Extraction Pipeline
The Font Handle Wrapper
LiteParse wraps PDFium’s low-level FPDF_FONT handle in a safe Rust type defined in crates/pdfium/src/font.rs. This Font struct provides methods to extract static properties once per text segment.
The base_name() method (lines 41-71) returns the PostScript font name with subset prefixes stripped, while font_type() (lines 73-84) returns an enum indicating whether the font is TrueType, Type1, or another format. The is_embedded() method (lines 87-90) checks whether the font data is embedded within the PDF file. These methods are invoked during the initial parsing of each text segment to capture the font’s identity and capabilities.
The TextItem Data Model
Extracted font metadata flows into the TextItem struct defined in crates/liteparse/src/types.rs (lines 13-45). This struct stores the font name, size, height, ascent, and descent values alongside the actual text content.
Crucially, TextItem includes a font_is_buggy boolean field. According to the source, this field is serialized only when true to minimize output size. When set, it signals that the segment contains glyphs from a font exhibiting encoding issues, allowing downstream processors to apply special handling rules.
Buggy Encoding Detection Heuristics
Font-Name Pattern Matching
LiteParse applies lightweight heuristics in crates/liteparse/src/extract.rs to identify fonts likely to have broken encodings. The is_buggy_font function (lines 501-516) checks for two specific patterns:
- TrueType subset fonts: Names starting with
"TT"or containing"+TT"(e.g.,TT1+Arial) - Type1 fonts: Names containing a six-character random prefix followed by an underscore (e.g.,
ABCDEF_FontName)
The function receives the font name and type, returning true if either pattern matches, prompting the system to treat the entire segment with caution.
Unicode Codepoint Validation
Beyond font names, LiteParse validates individual glyphs using the is_buggy_codepoint function (lines 518-520). This check flags:
- Control characters (Unicode ≤ 0x1F)
- Private-Use Area codepoints (0xE000–0xF8FF) commonly used by malformed PDF generators to map visible glyphs to non-standard Unicode regions
This validation runs once per segment during initialization and again for every subsequent glyph, ensuring that even a single problematic character flips the font_is_buggy flag for the entire text item.
Segment-Level Processing
The detection logic lives in SegmentBuilder within crates/liteparse/src/extract.rs. When processing begins, SegmentBuilder::start (lines 62-70) reads the font name and type; if the font is embedded and matches is_buggy_font, the segment’s internal font_is_buggy flag is set immediately (lines 68-70).
As the parser consumes each character via SegmentBuilder::push_char (lines 31-36), it checks the Unicode value using is_buggy_codepoint. If a control character or private-use codepoint appears, the flag is activated. When the segment is finally flushed to the output vector (lines 78-84), the current state of the flag is transferred to the TextItem struct, ensuring accurate propagation of font health status.
Accessing Font Metadata in Practice
Rust Implementation
The Rust API exposes the complete TextItem struct, allowing direct inspection of font metadata and the buggy flag:
use liteparse::{LiteParse, LiteParseConfig};
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let cfg = LiteParseConfig::default();
let parser = LiteParse::new(cfg);
let result = parser.parse("document.pdf").await?;
for page in result.pages {
for item in page.text_items {
if item.font_is_buggy {
println!(
"Warning: Buggy font '{}' (size {:.1}) with text: {}",
item.font_name.as_deref().unwrap_or("<unknown>"),
item.font_size.unwrap_or(0.0),
item.text
);
}
}
}
Ok(())
}
Node.js Bindings
The Node.js native bindings forward the full TextItem struct, including the boolean flag:
const { LiteParse } = require('liteparse');
(async () => {
const parser = new LiteParse();
const result = await parser.parse('document.pdf');
for (const page of result.pages) {
for (const item of page.text_items) {
if (item.font_is_buggy) {
console.log(
`Detected buggy encoding in font: ${item.font_name || 'unknown'}`
);
}
}
}
})();
Python Limitations
The high-level Python wrapper currently maps only a subset of TextItem fields (text, geometry, font name/size, and confidence), excluding the font_is_buggy flag. Developers requiring this metadata must either extend packages/python/liteparse/types.py to include the field or access the underlying native library directly via ctypes or cffi to read the full struct definition.
Summary
- Font metadata extraction occurs through a safe Rust wrapper around PDFium in
crates/pdfium/src/font.rs, providing access to PostScript names, font types, and embedding status. - Buggy font detection relies on dual heuristics in
crates/liteparse/src/extract.rs: pattern matching on font names (TrueType subsets and Type1 prefixes) and validation of Unicode codepoints (control characters and private-use areas). - The
font_is_buggyflag is set at the segment level inSegmentBuilderand stored in theTextItemstruct defined incrates/liteparse/src/types.rs, serialized only whentrue. - Language support varies: Rust and Node.js expose the full flag, while the Python wrapper requires modification to access this metadata.
Frequently Asked Questions
What triggers the font_is_buggy flag in LiteParse?
The flag activates when a font name matches specific corruption patterns—such as TrueType subset prefixes (TT or +TT) or Type1 random prefixes followed by an underscore—or when any glyph in the text segment contains a control character (≤ 0x1F) or falls within the Unicode Private-Use Area (0xE000–0xF8FF).
How does LiteParse extract clean font names from PDFs?
The library calls base_name() on the Font wrapper in crates/pdfium/src/font.rs, which interacts with PDFium to retrieve the PostScript name and automatically strips subset prefixes. This method runs once per text segment and returns the underlying typeface name regardless of how the PDF generator subset the font.
Why is the font_is_buggy field serialized only when true?
The TextItem struct uses Serde attributes to skip serializing the field when it is false, reducing JSON output size for the majority of text items that use standard, well-formed fonts. This optimization keeps responses compact while preserving critical metadata for the minority of segments requiring special handling.
Is buggy font detection available in all LiteParse language bindings?
No. The Rust core and Node.js bindings expose the font_is_buggy boolean directly in their TextItem implementations. The Python wrapper currently omits this field, requiring developers to either patch the type definitions or interface directly with the native library to access font health indicators.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →