What Is the Purpose and Effect of the preserve_very_small_text Configuration Option in LiteParse?
The preserve_very_small_text configuration option in LiteParse is a Boolean flag that, when set to true, disables the parser's internal filter and retains text fragments smaller than the hard-coded MIN_SMALL_TEXT_HEIGHT threshold in the final output.
When parsing PDFs with the open-source run-llama/liteparse library, most users want decorative glyphs, footnote markers, and noisy tiny symbols removed automatically. The preserve_very_small_text configuration option in LiteParse controls exactly this behavior, deciding whether sub-threshold text items survive the extraction pipeline. By default, the flag is false, which means the parser silently drops very small text unless you explicitly opt in.
Purpose of the preserve_very_small_text Configuration Option in LiteParse
preserve_very_small_text is a Boolean field on the LiteParseConfig struct declared in crates/liteparse/src/config.rs. It determines whether the parser keeps text items whose visual height or font size falls below an internal noise threshold.
According to the run-llama/liteparse source code, the flag is defined with a concise documentation comment:
/// Keep very small text that would normally be filtered out.
pub preserve_very_small_text: bool,
When this value is left at its default of false, LiteParse treats tiny glyphs as noise and excludes them from the document model. Setting it to true disables that cleanup step entirely.
Effect of the preserve_very_small_text Configuration Option on PDF Parsing
Enabling the preserve_very_small_text configuration option in LiteParse alters the extraction pipeline in three specific ways:
-
Skips the small-text-removal step. The parser normally drops items whose height is below the internal
MIN_SMALL_TEXT_HEIGHTconstant. Whenpreserve_very_small_textistrue, this filter is bypassed entirely. -
Propagates tiny items through layout projection and OCR merge. The retained fragments continue through the layout-projection and OCR-merge stages. They appear in the final JSON or plain-text export exactly as they were extracted from the PDF.
-
Increases output size and affects downstream processing. Because formerly ignored glyphs now become indexed tokens, downstream tasks such as text search, summarization, or embedding generation may ingest footnote markers, separator dots, and other decorative symbols that standard pipelines usually ignore.
Source Code Implementation Across LiteParse Bindings
The preserve_very_small_text flag is implemented consistently across the Rust core and all official language bindings.
Rust Core Declaration
The canonical definition lives in crates/liteparse/src/config.rs, where the field sits on the main LiteParseConfig struct alongside other parsing parameters.
CLI Mapping in main.rs
The LiteParse CLI exposes the feature through the --preserve-small-text flag. In crates/liteparse/src/main.rs, the argument is forwarded directly into the config struct:
preserve_very_small_text: cmd.preserve_small_text,
This one-line mapping ensures that command-line users can toggle the filter without editing configuration files.
Node.js, Python, and WebAssembly Bindings
All official wrappers surface the same Boolean:
- Node.js (N-API) —
LiteParseOptions.preserve_very_small_textis mapped to the Rust config incrates/liteparse-napi/src/types.rs. - Python — The constructor accepts
preserve_very_small_text=Trueand forwards it to the Rust core incrates/liteparse-python/src/lib.rs. - WebAssembly — The JavaScript wrapper includes
preserve_very_small_textas an optional field that maps to the Rust config incrates/liteparse-wasm/src/lib.rs.
How to Enable preserve_very_small_text in Code and CLI
You can activate the option through the Node.js, Python, or CLI interfaces.
Node.js Example
// Node.js – enable tiny‑text preservation
import { LiteParse } from "liteparse";
const parser = new LiteParse({
preserve_very_small_text: true, // keep every glyph, no matter how small
ocr_enabled: false,
});
await parser.parseFile("example.pdf");
console.log(parser.getResult());
Python Example
# Python – turn on the flag
from liteparse import LiteParse
parser = LiteParse(preserve_very_small_text=True, ocr_enabled=False)
parser.parse("example.pdf")
print(parser.result())
CLI Example
# CLI – pass the flag on the command line
liteparse --preserve-small-text --ocr-enabled=false input.pdf > out.json
Summary
- The
preserve_very_small_textconfiguration option in LiteParse is a Boolean flag defined incrates/liteparse/src/config.rsthat defaults tofalse. - When enabled, it skips the internal
MIN_SMALL_TEXT_HEIGHTfilter, allowing tiny glyphs to survive the layout-projection and OCR-merge stages. - The flag is exposed uniformly in the Rust core, CLI (
--preserve-small-text), Node.js, Python, and WebAssembly bindings. - Turning it on increases output fidelity but may introduce noise and enlarge the document model passed to downstream applications.
Frequently Asked Questions
What is the default value of preserve_very_small_text in LiteParse?
The default value is false. Unless you explicitly set the flag to true via code or the --preserve-small-text CLI option, LiteParse automatically removes text items whose height falls below the internal MIN_SMALL_TEXT_HEIGHT threshold.
Which LiteParse interfaces support the preserve_very_small_text option?
All major interfaces support it. The Rust core defines the flag in crates/liteparse/src/config.rs, the CLI accepts --preserve-small-text in crates/liteparse/src/main.rs, and the Node.js, Python, and WebAssembly bindings each expose the same Boolean field in their respective wrapper crates.
Does enabling preserve_very_small_text affect OCR results?
Yes. When the flag is true, tiny text fragments that would normally be discarded are propagated through the OCR-merge stage. This means even minuscule glyphs detected by OCR or extracted directly from the PDF will appear in the final output, potentially adding noise to OCR-dependent workflows.
Why would I want to keep very small text in a parsed PDF?
You should enable preserve_very_small_text when your pipeline requires absolute fidelity to the original document. Legal filings, academic papers with subscript citations, and financial reports with tiny footnote symbols are cases where dropping small glyphs could delete meaningful content rather than mere decoration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →