How to Preserve Very Small Text in LiteParse: Complete Configuration Guide

To preserve very small text in LiteParse output, set the min_font_size configuration option to 0.0 in your LiteParseConfig, ensuring the parser emits every character regardless of point size.

LiteParse is a Rust-based PDF text extraction library from the run-llama/liteparse repository that preserves the original font size of every character reported by PDFium. While the default behavior retains all text items, understanding the configuration options in the extraction pipeline ensures that tiny glyphs—such as footnote markers, superscripts, and fine print—survive the parsing process intact.

Understanding How LiteParse Handles Font Sizes

LiteParse stores font size metadata in the TextItem struct defined in crates/liteparse/src/types.rs. This structure carries the font_size field through the entire extraction pipeline, from raw PDF parsing in crates/liteparse/src/extract.rs through to the final output serializers in crates/liteparse/src/output/json.rs and crates/liteparse/src/output/text.rs.

By default, the parser does not discard items based on their point size. The only algorithmic consideration for font size occurs during layout reconstruction in crates/liteparse/src/projection.rs, where spatial grouping uses font-relative thresholds to determine line breaks and word boundaries. These thresholds are proportional to the current font size, meaning a 4 pt character generates a smaller tolerance than a 12 pt character, but the text itself is never filtered out.

Configuration Options for Small Text Preservation

The LiteParseConfig struct in crates/liteparse/src/config.rs provides explicit controls for handling microscopic text. Three specific settings affect whether tiny glyphs appear in your output:

  • min_font_size – An optional filter that removes TextItem instances below a specified point size. When set to 0.0 or omitted, all characters are preserved.
  • skip_invisible_text – Defaults to true and skips characters with PDF render mode 3 (invisible), but this filters based on rendering flags, not point size.
  • max_inline_gap – Defaults to approximately 15 pt and determines when characters separate into different words; this affects layout grouping but not preservation.

Disabling the Minimum Font Size Filter

To guarantee that even 1 pt text survives extraction, explicitly configure the minimum font size filter. In your configuration file or programmatic setup, set:

[filters]
min_font_size = 0.0

When this value is 0.0, the filter is effectively bypassed, allowing the extract.rs pipeline to pass all TextItem instances to the output stage regardless of how small the font_size value is.

Understanding Layout Projection Thresholds

The projection.rs module performs spatial analysis to reconstruct reading order. It calculates gap tolerances using the formula 0.5 × font_size for line-break detection and similar proportions for inline gaps. While very small text results in proportionally tiny thresholds, this behavior preserves the spatial relationship of fine details rather than discarding them. If you observe unintended line breaks in microscopic text, you can lower the max_inline_gap value, but this adjusts grouping logic without affecting text preservation.

Step-by-Step Implementation

Follow these steps to configure LiteParse for maximum text retention across all supported language bindings.

Rust Implementation

When using the library directly, instantiate LiteParse with a custom configuration:

use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.filters.min_font_size = Some(0.0);  // Preserve all text sizes
    
    let parser = LiteParse::new(cfg);
    let result = parser.parse_path("document.pdf")?;
    
    println!("{}", serde_json::to_string_pretty(&result)?);
    Ok(())
}

Node.js Implementation

The JavaScript wrapper in packages/node/src/lib.ts exposes the same configuration via camelCase options:

import { LiteParse } from "liteparse";

const parser = new LiteParse({
  filters: {
    minFontSize: 0  // Keep all glyphs including tiny text
  }
});

const output = await parser.parseFile("document.pdf");
console.log(JSON.stringify(output, null, 2));

Python Implementation

The Python wrapper mirrors the Rust configuration structure:

from liteparse import LiteParse, LiteParseConfig

cfg = LiteParseConfig()
cfg.filters.min_font_size = 0.0  # Ensure very small text is preserved

parser = LiteParse(cfg)
result = parser.parse_path("document.pdf")
print(result.json(indent=2))

Verifying Small Text in Output

After parsing, inspect the generated JSON structure to confirm preservation. Each entry in the text_items array contains a font_size field representing the exact point size reported by PDFium:

{
  "page_number": 1,
  "text_items": [
    {
      "text": "¹",
      "font_name": "Helvetica",
      "font_size": 4.0,
      "x": 150.2,
      "y": 720.5
    },
    {
      "text": "Micro-label",
      "font_name": "Arial",
      "font_size": 2.5,
      "x": 45.0,
      "y": 300.1
    }
  ]
}

If the font_size values reflect the microscopic sizes present in your source PDF (e.g., 2.5 or 4.0), the configuration is correctly preserving very small text in LiteParse output.

Summary

  • LiteParse stores font sizes in the TextItem struct (types.rs) and carries them through extract.rs to the output modules.
  • Set min_font_size to 0.0 in LiteParseConfig (config.rs) to disable size-based filtering and preserve all glyphs.
  • projection.rs uses font size only for spatial tolerance calculations during layout reconstruction, never for content removal.
  • Verify preservation by checking the font_size field in JSON output from output/json.rs.

Frequently Asked Questions

Does LiteParse filter text based on point size by default?

No. By default, LiteParse does not filter text based on point size. The only size-related logic appears in crates/liteparse/src/projection.rs, where font size helps calculate spatial thresholds for grouping characters into lines. To explicitly filter by size, you must opt-in by setting the min_font_size configuration option.

What is the minimum font size LiteParse can detect?

LiteParse can detect and preserve any font size that PDFium reports, including fractional point sizes below 1 pt. As long as min_font_size is set to 0.0 or omitted from your LiteParseConfig, the library will emit TextItem instances for all glyphs regardless of how small they are.

Why does my small text appear on separate lines?

This occurs in the layout reconstruction phase (projection.rs). The algorithm uses proportional thresholds based on font height to detect line breaks. Very small text generates tiny vertical tolerances, which can cause characters to split across lines if the PDF contains slight vertical variations. Adjust the max_inline_gap configuration to fine-tune grouping behavior without losing the text itself.

How do I configure LiteParse in a production environment?

Pass a configuration file to the CLI or instantiate LiteParseConfig programmatically. For the CLI (crates/liteparse/src/main.rs), use --config config.toml with min_font_size = 0.0 under the [filters] section. For library usage in Rust, Node.js (packages/node/src/lib.ts), or Python (packages/python/liteparse/parser.py), set the corresponding field on the configuration object before calling the parser.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →