# How to Preserve Very Small Text in LiteParse: Complete Configuration Guide

> Preserve very small text in LiteParse output by setting min_font_size to 0.0. Our guide shows you how to configure LiteParse for complete text extraction.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**To preserve very small text in LiteParse output, set the `min_font_size` configuration option to `0.0` in your `LiteParseConfig`, ensuring the parser emits every character regardless of point size.**

LiteParse is a Rust-based PDF text extraction library from the `run-llama/liteparse` repository that preserves the original font size of every character reported by PDFium. While the default behavior retains all text items, understanding the configuration options in the extraction pipeline ensures that tiny glyphs—such as footnote markers, superscripts, and fine print—survive the parsing process intact.

## Understanding How LiteParse Handles Font Sizes

LiteParse stores font size metadata in the `TextItem` struct defined in **[`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs)**. This structure carries the `font_size` field through the entire extraction pipeline, from raw PDF parsing in **[`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs)** through to the final output serializers in **[`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs)** and **[`crates/liteparse/src/output/text.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/text.rs)**.

By default, the parser does **not** discard items based on their point size. The only algorithmic consideration for font size occurs during layout reconstruction in **[`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)**, where spatial grouping uses font-relative thresholds to determine line breaks and word boundaries. These thresholds are proportional to the current font size, meaning a 4 pt character generates a smaller tolerance than a 12 pt character, but the text itself is never filtered out.

## Configuration Options for Small Text Preservation

The `LiteParseConfig` struct in **[`crates/liteparse/src/config.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/config.rs)** provides explicit controls for handling microscopic text. Three specific settings affect whether tiny glyphs appear in your output:

- **`min_font_size`** – An optional filter that removes `TextItem` instances below a specified point size. When set to `0.0` or omitted, all characters are preserved.
- **`skip_invisible_text`** – Defaults to `true` and skips characters with PDF render mode 3 (invisible), but this filters based on rendering flags, not point size.
- **`max_inline_gap`** – Defaults to approximately 15 pt and determines when characters separate into different words; this affects layout grouping but not preservation.

### Disabling the Minimum Font Size Filter

To guarantee that even 1 pt text survives extraction, explicitly configure the minimum font size filter. In your configuration file or programmatic setup, set:

```toml
[filters]
min_font_size = 0.0

```

When this value is `0.0`, the filter is effectively bypassed, allowing the [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs) pipeline to pass all `TextItem` instances to the output stage regardless of how small the `font_size` value is.

### Understanding Layout Projection Thresholds

The **[`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs)** module performs spatial analysis to reconstruct reading order. It calculates gap tolerances using the formula `0.5 × font_size` for line-break detection and similar proportions for inline gaps. While very small text results in proportionally tiny thresholds, this behavior preserves the spatial relationship of fine details rather than discarding them. If you observe unintended line breaks in microscopic text, you can lower the `max_inline_gap` value, but this adjusts grouping logic without affecting text preservation.

## Step-by-Step Implementation

Follow these steps to configure LiteParse for maximum text retention across all supported language bindings.

### Rust Implementation

When using the library directly, instantiate `LiteParse` with a custom configuration:

```rust
use liteparse::LiteParse;
use liteparse::config::LiteParseConfig;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut cfg = LiteParseConfig::default();
    cfg.filters.min_font_size = Some(0.0);  // Preserve all text sizes
    
    let parser = LiteParse::new(cfg);
    let result = parser.parse_path("document.pdf")?;
    
    println!("{}", serde_json::to_string_pretty(&result)?);
    Ok(())
}

```

### Node.js Implementation

The JavaScript wrapper in **[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)** exposes the same configuration via camelCase options:

```javascript
import { LiteParse } from "liteparse";

const parser = new LiteParse({
  filters: {
    minFontSize: 0  // Keep all glyphs including tiny text
  }
});

const output = await parser.parseFile("document.pdf");
console.log(JSON.stringify(output, null, 2));

```

### Python Implementation

The Python wrapper mirrors the Rust configuration structure:

```python
from liteparse import LiteParse, LiteParseConfig

cfg = LiteParseConfig()
cfg.filters.min_font_size = 0.0  # Ensure very small text is preserved

parser = LiteParse(cfg)
result = parser.parse_path("document.pdf")
print(result.json(indent=2))

```

## Verifying Small Text in Output

After parsing, inspect the generated JSON structure to confirm preservation. Each entry in the `text_items` array contains a `font_size` field representing the exact point size reported by PDFium:

```json
{
  "page_number": 1,
  "text_items": [
    {
      "text": "¹",
      "font_name": "Helvetica",
      "font_size": 4.0,
      "x": 150.2,
      "y": 720.5
    },
    {
      "text": "Micro-label",
      "font_name": "Arial",
      "font_size": 2.5,
      "x": 45.0,
      "y": 300.1
    }
  ]
}

```

If the `font_size` values reflect the microscopic sizes present in your source PDF (e.g., `2.5` or `4.0`), the configuration is correctly preserving very small text in LiteParse output.

## Summary

- **LiteParse stores font sizes** in the `TextItem` struct (**[`types.rs`](https://github.com/run-llama/liteparse/blob/main/types.rs)**) and carries them through **[`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs)** to the output modules.
- **Set `min_font_size` to `0.0`** in `LiteParseConfig` (**[`config.rs`](https://github.com/run-llama/liteparse/blob/main/config.rs)**) to disable size-based filtering and preserve all glyphs.
- **[`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs)** uses font size only for spatial tolerance calculations during layout reconstruction, never for content removal.
- **Verify preservation** by checking the `font_size` field in JSON output from **[`output/json.rs`](https://github.com/run-llama/liteparse/blob/main/output/json.rs)**.

## Frequently Asked Questions

### Does LiteParse filter text based on point size by default?

No. By default, LiteParse does not filter text based on point size. The only size-related logic appears in **[`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs)**, where font size helps calculate spatial thresholds for grouping characters into lines. To explicitly filter by size, you must opt-in by setting the `min_font_size` configuration option.

### What is the minimum font size LiteParse can detect?

LiteParse can detect and preserve any font size that PDFium reports, including fractional point sizes below 1 pt. As long as `min_font_size` is set to `0.0` or omitted from your `LiteParseConfig`, the library will emit `TextItem` instances for all glyphs regardless of how small they are.

### Why does my small text appear on separate lines?

This occurs in the layout reconstruction phase (**[`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs)**). The algorithm uses proportional thresholds based on font height to detect line breaks. Very small text generates tiny vertical tolerances, which can cause characters to split across lines if the PDF contains slight vertical variations. Adjust the `max_inline_gap` configuration to fine-tune grouping behavior without losing the text itself.

### How do I configure LiteParse in a production environment?

Pass a configuration file to the CLI or instantiate `LiteParseConfig` programmatically. For the CLI (**[`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs)**), use `--config config.toml` with `min_font_size = 0.0` under the `[filters]` section. For library usage in Rust, Node.js (**[`packages/node/src/lib.ts`](https://github.com/run-llama/liteparse/blob/main/packages/node/src/lib.ts)**), or Python (**[`packages/python/liteparse/parser.py`](https://github.com/run-llama/liteparse/blob/main/packages/python/liteparse/parser.py)**), set the corresponding field on the configuration object before calling the parser.