Debugging LiteParse Parsing Issues: A Step-by-Step Source Code Guide

You can debug LiteParse parsing issues by tracing data through eight well-defined pipeline stages—from raw PDF extraction in extract.rs to layout reconstruction in projection.rs—using CLI debug logs, intermediate ProjectedTextItem dumps, and the integration test suite as a reproducible baseline.

Follow the flow of data through run-llama/liteparse to pinpoint exactly where layout reconstruction diverges from expectations. The crate processes documents through a strict pipeline, and each stage exposes distinct inspection points in the source code. Understanding these stages is the fastest way to resolve misaligned columns, missing paragraphs, or incorrect rotations.

Understand the LiteParse Pipeline

LiteParse processes PDFs through a strict sequence of discrete transformations. The journey starts at the CLI entry point in crates/liteparse/src/main.rs and the public LiteParse struct in crates/liteparse/src/lib.rs, then moves through raw extraction, rotation normalization, anchor mapping, flowing-text classification, and optional OCR merging. Every stage returns a Result type defined in crates/liteparse/src/error.rs, so you can isolate failures by examining the module where the error originates.

Debug LiteParse Parsing Issues by Pipeline Stage

Verify the CLI or API Entry Point

Start by confirming that your LiteParseConfig is initialized correctly. In crates/liteparse/src/main.rs, the CLI creates the configuration and hands it to parser::parse, while crates/liteparse/src/lib.rs exposes the same public API for programmatic use. If the output is unexpectedly empty, ensure that your PDF path and flags are reaching the internal parser.

Inspect Raw Text Extraction

Before layout reconstruction begins, PDFium returns a flat list of TextItem objects in crates/liteparse/src/extract.rs. Verify that the raw list contains the expected number of items, correct coordinates, and text strings by checking the extract_page implementation. If the raw extraction is missing text, the problem is upstream of layout logic and likely stems from the PDFium binding or the input file itself.

Check Rotation Normalization

Rotated text is normalized before layout reconstruction in crates/liteparse/src/projection.rs. Look at canonical_rotation, which rounds angles to 0/90/180/270°, and handle_rotation_reading_order, which groups items by rotation and re-orders them. When text appears out of order on rotated pages, inspect these functions around lines 89 and 113 to confirm that rotation values are being bucketed and transformed correctly.

Audit Anchor Detection and Filtering

The engine builds left, right, and center anchor maps to snap boxes together, defined in crates/liteparse/src/projection.rs. Two filters are common culprits when columns are mis-detected: delta_min_filter (removes isolated anchors) around line 549, and intercept_filter (drops anchors that are visually crossed) around line 995. If your table columns are merging or splitting incorrectly, add debug prints before and after these filters to inspect anchor map sizes.

Tune Flowing-Text Constants

A set of FLOWING_* constants at the top of crates/liteparse/src/projection.rs controls when a block is treated as flowing prose rather than a table. FLOWING_MAX_TOTAL_ANCHORS and FLOWING_WIDE_LINE_RATIO (currently 0.5) determine paragraph boundaries. If the parser mistakenly splits a paragraph into separate lines, try increasing FLOWING_WIDE_LINE_RATIO or decreasing FLOWING_MIN_LINE_ITEMS in a local copy and re-running the test binary.

Validate OCR Result Merging

When OCR is enabled, crates/liteparse/src/ocr_merge.rs combines native PDF text with OCR-generated text. Verify that the merged flag is set correctly and that item.rotation is cleared after merging. If you see duplicated or misaligned text, print the OCR-only items before merging to confirm they align to the same x/y range as the native items.

Read Error Variants for Stage Identification

All stages propagate failures through crates/liteparse/src/error.rs, which provides detailed variants such as PdfiumError and OcrError. When a parsing step fails, the error message itself indicates the stage because each module returns the centralized LiteParseError type. Use the variant name to decide whether to investigate PDFium bindings, OCR workers, or layout projection logic.

Baseline Behavior with Integration Tests

The integration test in crates/liteparse/tests/integration_test.rs runs a full parse on a known PDF and compares the JSON output. Use this file as a sanity check, or replace the fixture PDF with your problematic file and observe which assertion fails. This test is the fastest way to create a minimal reproducible case that exercises the entire pipeline from crates/liteparse/src/parser.rs.

Practical Debugging Workflow

  1. Run the CLI with maximum verbosity. Add --log=debug to forward the flag to env_logger.

    liteparse mydoc.pdf --output json --log=debug > out.json

    The log output includes the number of raw items extracted in extract.rs, the rotation groups detected in handle_rotation_reading_order, and anchor map sizes before and after delta_min_filter and intercept_filter.

  2. Inspect the intermediate ProjectedTextItem vector. In Rust, temporarily add a debug dump in parser.rs after the call to project_items:

    let mut projected = project_items(raw_items, &config)?;
    eprintln!("PROJECTED ITEMS ({})", projected.len());
    for (i, item) in projected.iter().enumerate() {
        eprintln!("{}: {:?} – rot:{:?}", i, item.item, item.item.rotation);
    }

    This tells you whether rotation has been cleared and whether items have sensible coordinates before layout reconstruction begins.

  3. Validate OCR merging. If you use OCR, print the OCR-only items before they reach the merge stage:

    let ocr_items = run_ocr(&rendered_pages)?;
    eprintln!("OCR items: {:?}", ocr_items);

    Confirm that the OCR text aligns to the same x/y range as the native items in crates/liteparse/src/ocr_merge.rs.

  4. Tune the flowing-text thresholds. If whole paragraphs appear as separate lines, adjust the constants in crates/liteparse/src/projection.rs.

    Try increasing FLOWING_WIDE_LINE_RATIO (currently 0.5) or decreasing FLOWING_MIN_LINE_ITEMS. Re-run the test binary after each change and compare the output.

  5. Re-run the integration test with a custom PDF. Swap the fixture in crates/liteparse/tests/integration_test.rs for your problematic file and observe which assertion fails. The test prints the parsed JSON, which you can diff against the expected output to identify the exact stage where output diverges.

Debug LiteParse from Node.js, Python, and the CLI

Node.js Wrapper

Enable debugging and inspect the parsed structure before it reaches your application:

import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse({ logLevel: "debug", enableOcr: false });
  const result = await parser.parse("sample.pdf");
  console.log(JSON.stringify(result, null, 2));

  // Dump raw projected items by adding this method in packages/node/src/native.ts:
  // native.dumpProjectedItems();
})();

Python Wrapper

Capture the parser output and hook into the native layer for intermediate state:

from liteparse import LiteParse

parser = LiteParse(log_level="debug", ocr=False)
result = parser.parse("sample.pdf")
print(result.json(indent=2))

# To dump intermediate ProjectedTextItems, add to liteparse/parser.py:

#   from ._native import _dump_projected_items

#   _dump_projected_items()

CLI Rotation Override

Force a specific rotation handling path to see how the normalizer behaves:

liteparse rotated.pdf --disable-ocr --log=debug --rotate-force=90

This helps you verify how canonical_rotation and handle_rotation_reading_order in crates/liteparse/src/projection.rs process rotated labels.

Summary

Frequently Asked Questions

How do I know which LiteParse stage is causing a parsing error?

Every stage returns a Result<…, LiteParseError> defined in crates/liteparse/src/error.rs. The error variant—such as PdfiumError or OcrError—tells you exactly which module failed, so you can inspect extract.rs, ocr_merge.rs, or projection.rs accordingly.

Why are my table columns merging or splitting incorrectly in LiteParse?

Incorrect column detection usually traces back to the anchor filters in crates/liteparse/src/projection.rs. delta_min_filter removes isolated anchors around line 549, and intercept_filter drops visually crossed anchors around line 995. Add debug prints before and after these filters to inspect anchor map sizes.

Inspect crates/liteparse/src/ocr_merge.rs and verify that the merged flag is set correctly and that item.rotation is cleared after merging. You can also print the OCR-only items before merging to confirm their x/y coordinates align with the native text extracted from crates/liteparse/src/extract.rs.

What is the fastest way to reproduce a LiteParse bug for debugging?

Replace the fixture PDF in crates/liteparse/tests/integration_test.rs with your problematic file and run the integration test. This end-to-end test compares the parsed JSON against expected output, so the first failing assertion points to the exact layout stage that diverges from expectations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →