How to Preserve Very Small Text in LiteParse: Complete Configuration Guide
To preserve very small text in LiteParse output, set the min_font_size configuration option to 0.0 in your LiteParseConfig, ensuring the parser emits every character regardless of point size.
LiteParse is a Rust-based PDF text extraction library from the run-llama/liteparse repository that preserves the original font size of every character reported by PDFium. While the default behavior retains all text items, understanding the configuration options in the extraction pipeline ensures that tiny glyphs—such as footnote markers, superscripts, and fine print—survive the parsing process intact.
Understanding How LiteParse Handles Font Sizes
LiteParse stores font size metadata in the TextItem struct defined in crates/liteparse/src/types.rs. This structure carries the font_size field through the entire extraction pipeline, from raw PDF parsing in crates/liteparse/src/extract.rs through to the final output serializers in crates/liteparse/src/output/json.rs and crates/liteparse/src/output/text.rs.
By default, the parser does not discard items based on their point size. The only algorithmic consideration for font size occurs during layout reconstruction in crates/liteparse/src/projection.rs, where spatial grouping uses font-relative thresholds to determine line breaks and word boundaries. These thresholds are proportional to the current font size, meaning a 4 pt character generates a smaller tolerance than a 12 pt character, but the text itself is never filtered out.
Configuration Options for Small Text Preservation
The LiteParseConfig struct in crates/liteparse/src/config.rs provides explicit controls for handling microscopic text. Three specific settings affect whether tiny glyphs appear in your output:
min_font_size– An optional filter that removesTextIteminstances below a specified point size. When set to0.0or omitted, all characters are preserved.skip_invisible_text– Defaults totrueand skips characters with PDF render mode 3 (invisible), but this filters based on rendering flags, not point size.max_inline_gap– Defaults to approximately 15 pt and determines when characters separate into different words; this affects layout grouping but not preservation.
Disabling the Minimum Font Size Filter
To guarantee that even 1 pt text survives extraction, explicitly configure the minimum font size filter. In your configuration file or programmatic setup, set:
[filters]
min_font_size = 0.0
When this value is 0.0, the filter is effectively bypassed, allowing the extract.rs pipeline to pass all TextItem instances to the output stage regardless of how small the font_size value is.
Understanding Layout Projection Thresholds
The projection.rs module performs spatial analysis to reconstruct reading order. It calculates gap tolerances using the formula 0.5 × font_size for line-break detection and similar proportions for inline gaps. While very small text results in proportionally tiny thresholds, this behavior preserves the spatial relationship of fine details rather than discarding them. If you observe unintended line breaks in microscopic text, you can lower the max_inline_gap value, but this adjusts grouping logic without affecting text preservation.
Step-by-Step Implementation
Follow these steps to configure LiteParse for maximum text retention across all supported language bindings.
Rust Implementation
When using the library directly, instantiate LiteParse with a custom configuration:
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let mut cfg = LiteParseConfig::default();
cfg.filters.min_font_size = Some(0.0); // Preserve all text sizes
let parser = LiteParse::new(cfg);
let result = parser.parse_path("document.pdf")?;
println!("{}", serde_json::to_string_pretty(&result)?);
Ok(())
}
Node.js Implementation
The JavaScript wrapper in packages/node/src/lib.ts exposes the same configuration via camelCase options:
import { LiteParse } from "liteparse";
const parser = new LiteParse({
filters: {
minFontSize: 0 // Keep all glyphs including tiny text
}
});
const output = await parser.parseFile("document.pdf");
console.log(JSON.stringify(output, null, 2));
Python Implementation
The Python wrapper mirrors the Rust configuration structure:
from liteparse import LiteParse, LiteParseConfig
cfg = LiteParseConfig()
cfg.filters.min_font_size = 0.0 # Ensure very small text is preserved
parser = LiteParse(cfg)
result = parser.parse_path("document.pdf")
print(result.json(indent=2))
Verifying Small Text in Output
After parsing, inspect the generated JSON structure to confirm preservation. Each entry in the text_items array contains a font_size field representing the exact point size reported by PDFium:
{
"page_number": 1,
"text_items": [
{
"text": "¹",
"font_name": "Helvetica",
"font_size": 4.0,
"x": 150.2,
"y": 720.5
},
{
"text": "Micro-label",
"font_name": "Arial",
"font_size": 2.5,
"x": 45.0,
"y": 300.1
}
]
}
If the font_size values reflect the microscopic sizes present in your source PDF (e.g., 2.5 or 4.0), the configuration is correctly preserving very small text in LiteParse output.
Summary
- LiteParse stores font sizes in the
TextItemstruct (types.rs) and carries them throughextract.rsto the output modules. - Set
min_font_sizeto0.0inLiteParseConfig(config.rs) to disable size-based filtering and preserve all glyphs. projection.rsuses font size only for spatial tolerance calculations during layout reconstruction, never for content removal.- Verify preservation by checking the
font_sizefield in JSON output fromoutput/json.rs.
Frequently Asked Questions
Does LiteParse filter text based on point size by default?
No. By default, LiteParse does not filter text based on point size. The only size-related logic appears in crates/liteparse/src/projection.rs, where font size helps calculate spatial thresholds for grouping characters into lines. To explicitly filter by size, you must opt-in by setting the min_font_size configuration option.
What is the minimum font size LiteParse can detect?
LiteParse can detect and preserve any font size that PDFium reports, including fractional point sizes below 1 pt. As long as min_font_size is set to 0.0 or omitted from your LiteParseConfig, the library will emit TextItem instances for all glyphs regardless of how small they are.
Why does my small text appear on separate lines?
This occurs in the layout reconstruction phase (projection.rs). The algorithm uses proportional thresholds based on font height to detect line breaks. Very small text generates tiny vertical tolerances, which can cause characters to split across lines if the PDF contains slight vertical variations. Adjust the max_inline_gap configuration to fine-tune grouping behavior without losing the text itself.
How do I configure LiteParse in a production environment?
Pass a configuration file to the CLI or instantiate LiteParseConfig programmatically. For the CLI (crates/liteparse/src/main.rs), use --config config.toml with min_font_size = 0.0 under the [filters] section. For library usage in Rust, Node.js (packages/node/src/lib.ts), or Python (packages/python/liteparse/parser.py), set the corresponding field on the configuration object before calling the parser.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →