# Understanding the Confidence Field in LiteParse's TextItem: Purpose and Calculation

> Explore the confidence field in LiteParse's TextItem. Learn its purpose and how the 0.0-1.0 score indicates trust in OCR text, with native PDF text scoring 1.0.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**The `confidence` field in LiteParse's `TextItem` tells downstream consumers how much to trust OCR-derived text, using a normalized 0.0–1.0 score where native PDF text defaults to `1.0` in JSON output.**

When processing documents with the [run-llama/liteparse](https://github.com/run-llama/liteparse) library, every extracted text fragment is represented as a `TextItem`. The **confidence field in LiteParse's `TextItem`** communicates the reliability of the extracted string, distinguishing between native PDF text and OCR-generated content. Understanding how this value is computed and normalized is critical for any pipeline that filters or weights text based on recognition accuracy.

## Purpose of the Confidence Field in LiteParse's `TextItem`

In [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs), the `TextItem` struct defines `confidence` with a comment describing it as an "OCR confidence score (0.0–1.0). None for native PDF text." This makes the field's purpose explicit: it is a trust signal for OCR output.

- A value of `1.0` means maximum certainty.
- Values between `0.0` and `1.0` indicate decreasing reliability.
- A value of `None` means the text was extracted directly from the PDF's internal text streams and was not produced by an OCR engine.

Downstream consumers—including JSON serializers, language bindings, and custom post-processing logic—use this metric to decide how aggressively to trust or discard a given text fragment.

## How LiteParse Determines the `confidence` Value

The `confidence` value follows a five-step pipeline that transforms raw OCR engine output into the final normalized score stored in each `TextItem`.

1. **Raw OCR engine output** — Each OCR implementation, whether Tesseract or an HTTP-based OCR server, returns a per-word confidence score on a `0–100` scale.
2. **Normalization to `0.0–1.0`** — In [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs), LiteParse divides the raw score by `100.0`. The source code contains the comment `// tesseract-rs returns confidence 0-100, normalize to 0-1`.
3. **Low-confidence filtering** — The Tesseract wrapper discards any word whose normalized confidence is less than or equal to `0.30`, matching the behavior of the original TypeScript implementation.
4. **Merging with native text** — When OCR results replace missing native segments, [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) stores the OCR word confidence in the `TextItem` and rounds the value to three decimal places.
5. **Serialization default** — When building JSON in [`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs), the serializer substitutes `Some(1.0)` whenever a `TextItem` has `confidence == None`, guaranteeing that every output object carries a numeric confidence field.

## JSON and Rust Examples

Every `TextItem` in the final JSON output is guaranteed to include a numeric `confidence` field because the serializer in [`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs) normalizes `None` to `1.0` for native text.

```json
{
  "pages": [
    {
      "page_number": 1,
      "text_items": [
        {
          "text": "Hello",
          "x": 72.0,
          "y": 100.5,
          "confidence": 1.0
        },
        {
          "text": "World",
          "x": 120.0,
          "y": 100.5,
          "confidence": 0.84
        }
      ]
    }
  ]
}

```

In this output, `"Hello"` originated from native PDF text and was defaulted to `1.0`, while `"World"` came from OCR and retained its normalized score of `0.84`.

When consuming the Rust API directly, you can inspect `item.confidence` to branch based on text origin:

```rust
use liteparse::LiteParse;

let parser = LiteParse::new("example.pdf");
let result = parser.parse().await?;

for page in result.pages {
    for item in page.text_items {
        match item.confidence {
            Some(conf) => println!("'{}' (conf: {:.2})", item.text, conf),
            None => println!("'{}' (native PDF text)", item.text),
        }
    }
}

```

Note that `None` only appears in the internal Rust structure; after JSON serialization the field is always present as a number.

## Summary

- The **confidence field in LiteParse's `TextItem`** signals how trustworthy OCR-derived text is on a normalized `0.0–1.0` scale.
- **Native PDF text** sets `confidence` to `None` internally, but JSON output normalizes this to `1.0` in [`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs).
- **OCR scores** are normalized from `0–100` to `0.0–1.0` in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs).
- Results at or below **0.30** confidence are filtered out by the Tesseract wrapper.
- During merge in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs), the score is rounded to three decimal places before being stored.

## Frequently Asked Questions

### What does a `confidence` value of `1.0` mean in LiteParse?

A `confidence` value of `1.0` indicates maximum certainty. In the JSON output, this value is assigned to all native PDF text because it was extracted directly from the document's internal text streams rather than through OCR. For OCR-derived text, a score of `1.0` would mean the engine reported absolute certainty on its 0–100 scale.

### Why is the `confidence` field `None` for some `TextItem` objects?

Inside the Rust API, `confidence` is `None` whenever the text was extracted natively from the PDF without OCR. The field is only populated when the text originates from an OCR engine. The JSON serializer later replaces `None` with `1.0` so that downstream systems always receive a numeric value.

### How does LiteParse handle low-confidence OCR words?

LiteParse's Tesseract wrapper in [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) drops any word whose normalized confidence is less than or equal to `0.30`. This prevents noisy or unrecognizable glyphs from entering the parse result.

### Which source files control the `confidence` calculation in LiteParse?

The calculation and handling are spread across four key files: [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) defines the field semantics; [`crates/liteparse/src/ocr/tesseract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr/tesseract.rs) normalizes raw OCR scores and filters low-confidence words; [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) merges the OCR words into the parse tree and rounds the value; and [`crates/liteparse/src/output/json.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/output/json.rs) ensures the JSON output always contains a numeric confidence value.