Understanding the Confidence Field in LiteParse's TextItem: Purpose and Calculation

The confidence field in LiteParse's TextItem tells downstream consumers how much to trust OCR-derived text, using a normalized 0.0–1.0 score where native PDF text defaults to 1.0 in JSON output.

When processing documents with the run-llama/liteparse library, every extracted text fragment is represented as a TextItem. The confidence field in LiteParse's TextItem communicates the reliability of the extracted string, distinguishing between native PDF text and OCR-generated content. Understanding how this value is computed and normalized is critical for any pipeline that filters or weights text based on recognition accuracy.

Purpose of the Confidence Field in LiteParse's TextItem

In crates/liteparse/src/types.rs, the TextItem struct defines confidence with a comment describing it as an "OCR confidence score (0.0–1.0). None for native PDF text." This makes the field's purpose explicit: it is a trust signal for OCR output.

  • A value of 1.0 means maximum certainty.
  • Values between 0.0 and 1.0 indicate decreasing reliability.
  • A value of None means the text was extracted directly from the PDF's internal text streams and was not produced by an OCR engine.

Downstream consumers—including JSON serializers, language bindings, and custom post-processing logic—use this metric to decide how aggressively to trust or discard a given text fragment.

How LiteParse Determines the confidence Value

The confidence value follows a five-step pipeline that transforms raw OCR engine output into the final normalized score stored in each TextItem.

  1. Raw OCR engine output — Each OCR implementation, whether Tesseract or an HTTP-based OCR server, returns a per-word confidence score on a 0–100 scale.
  2. Normalization to 0.0–1.0 — In crates/liteparse/src/ocr/tesseract.rs, LiteParse divides the raw score by 100.0. The source code contains the comment // tesseract-rs returns confidence 0-100, normalize to 0-1.
  3. Low-confidence filtering — The Tesseract wrapper discards any word whose normalized confidence is less than or equal to 0.30, matching the behavior of the original TypeScript implementation.
  4. Merging with native text — When OCR results replace missing native segments, crates/liteparse/src/ocr_merge.rs stores the OCR word confidence in the TextItem and rounds the value to three decimal places.
  5. Serialization default — When building JSON in crates/liteparse/src/output/json.rs, the serializer substitutes Some(1.0) whenever a TextItem has confidence == None, guaranteeing that every output object carries a numeric confidence field.

JSON and Rust Examples

Every TextItem in the final JSON output is guaranteed to include a numeric confidence field because the serializer in crates/liteparse/src/output/json.rs normalizes None to 1.0 for native text.

{
  "pages": [
    {
      "page_number": 1,
      "text_items": [
        {
          "text": "Hello",
          "x": 72.0,
          "y": 100.5,
          "confidence": 1.0
        },
        {
          "text": "World",
          "x": 120.0,
          "y": 100.5,
          "confidence": 0.84
        }
      ]
    }
  ]
}

In this output, "Hello" originated from native PDF text and was defaulted to 1.0, while "World" came from OCR and retained its normalized score of 0.84.

When consuming the Rust API directly, you can inspect item.confidence to branch based on text origin:

use liteparse::LiteParse;

let parser = LiteParse::new("example.pdf");
let result = parser.parse().await?;

for page in result.pages {
    for item in page.text_items {
        match item.confidence {
            Some(conf) => println!("'{}' (conf: {:.2})", item.text, conf),
            None => println!("'{}' (native PDF text)", item.text),
        }
    }
}

Note that None only appears in the internal Rust structure; after JSON serialization the field is always present as a number.

Summary

  • The confidence field in LiteParse's TextItem signals how trustworthy OCR-derived text is on a normalized 0.0–1.0 scale.
  • Native PDF text sets confidence to None internally, but JSON output normalizes this to 1.0 in crates/liteparse/src/output/json.rs.
  • OCR scores are normalized from 0–100 to 0.0–1.0 in crates/liteparse/src/ocr/tesseract.rs.
  • Results at or below 0.30 confidence are filtered out by the Tesseract wrapper.
  • During merge in crates/liteparse/src/ocr_merge.rs, the score is rounded to three decimal places before being stored.

Frequently Asked Questions

What does a confidence value of 1.0 mean in LiteParse?

A confidence value of 1.0 indicates maximum certainty. In the JSON output, this value is assigned to all native PDF text because it was extracted directly from the document's internal text streams rather than through OCR. For OCR-derived text, a score of 1.0 would mean the engine reported absolute certainty on its 0–100 scale.

Why is the confidence field None for some TextItem objects?

Inside the Rust API, confidence is None whenever the text was extracted natively from the PDF without OCR. The field is only populated when the text originates from an OCR engine. The JSON serializer later replaces None with 1.0 so that downstream systems always receive a numeric value.

How does LiteParse handle low-confidence OCR words?

LiteParse's Tesseract wrapper in crates/liteparse/src/ocr/tesseract.rs drops any word whose normalized confidence is less than or equal to 0.30. This prevents noisy or unrecognizable glyphs from entering the parse result.

Which source files control the confidence calculation in LiteParse?

The calculation and handling are spread across four key files: crates/liteparse/src/types.rs defines the field semantics; crates/liteparse/src/ocr/tesseract.rs normalizes raw OCR scores and filters low-confidence words; crates/liteparse/src/ocr_merge.rs merges the OCR words into the parse tree and rounds the value; and crates/liteparse/src/output/json.rs ensures the JSON output always contains a numeric confidence value.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →