# Debugging LiteParse Parsing Issues: A Step-by-Step Source Code Guide

> Debug LiteParse parsing issues with this source code guide. Trace data through eight pipeline stages using CLI debug logs and intermediate dumps for effective troubleshooting.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: how-to-guide
- Published: 2026-06-07

---

**You can debug LiteParse parsing issues by tracing data through eight well-defined pipeline stages—from raw PDF extraction in [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs) to layout reconstruction in [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs)—using CLI debug logs, intermediate `ProjectedTextItem` dumps, and the integration test suite as a reproducible baseline.**

Follow the flow of data through `run-llama/liteparse` to pinpoint exactly where layout reconstruction diverges from expectations. The crate processes documents through a strict pipeline, and each stage exposes distinct inspection points in the source code. Understanding these stages is the fastest way to resolve misaligned columns, missing paragraphs, or incorrect rotations.

## Understand the LiteParse Pipeline

LiteParse processes PDFs through a strict sequence of discrete transformations. The journey starts at the CLI entry point in [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) and the public `LiteParse` struct in [`crates/liteparse/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/lib.rs), then moves through raw extraction, rotation normalization, anchor mapping, flowing-text classification, and optional OCR merging. Every stage returns a `Result` type defined in [`crates/liteparse/src/error.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/error.rs), so you can isolate failures by examining the module where the error originates.

## Debug LiteParse Parsing Issues by Pipeline Stage

### Verify the CLI or API Entry Point

Start by confirming that your **`LiteParseConfig`** is initialized correctly. In [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs), the CLI creates the configuration and hands it to **`parser::parse`**, while [`crates/liteparse/src/lib.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/lib.rs) exposes the same public API for programmatic use. If the output is unexpectedly empty, ensure that your PDF path and flags are reaching the internal parser.

### Inspect Raw Text Extraction

Before layout reconstruction begins, PDFium returns a flat list of **`TextItem`** objects in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs). Verify that the raw list contains the expected number of items, correct coordinates, and text strings by checking the **`extract_page`** implementation. If the raw extraction is missing text, the problem is upstream of layout logic and likely stems from the PDFium binding or the input file itself.

### Check Rotation Normalization

Rotated text is normalized before layout reconstruction in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs). Look at **`canonical_rotation`**, which rounds angles to 0/90/180/270°, and **`handle_rotation_reading_order`**, which groups items by rotation and re-orders them. When text appears out of order on rotated pages, inspect these functions around lines 89 and 113 to confirm that rotation values are being bucketed and transformed correctly.

### Audit Anchor Detection and Filtering

The engine builds left, right, and center anchor maps to snap boxes together, defined in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs). Two filters are common culprits when columns are mis-detected: **`delta_min_filter`** (removes isolated anchors) around line 549, and **`intercept_filter`** (drops anchors that are visually crossed) around line 995. If your table columns are merging or splitting incorrectly, add debug prints before and after these filters to inspect anchor map sizes.

### Tune Flowing-Text Constants

A set of **`FLOWING_*`** constants at the top of [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) controls when a block is treated as flowing prose rather than a table. **`FLOWING_MAX_TOTAL_ANCHORS`** and **`FLOWING_WIDE_LINE_RATIO`** (currently `0.5`) determine paragraph boundaries. If the parser mistakenly splits a paragraph into separate lines, try increasing `FLOWING_WIDE_LINE_RATIO` or decreasing **`FLOWING_MIN_LINE_ITEMS`** in a local copy and re-running the test binary.

### Validate OCR Result Merging

When OCR is enabled, [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) combines native PDF text with OCR-generated text. Verify that the **`merged`** flag is set correctly and that **`item.rotation`** is cleared after merging. If you see duplicated or misaligned text, print the OCR-only items before merging to confirm they align to the same `x/y` range as the native items.

### Read Error Variants for Stage Identification

All stages propagate failures through [`crates/liteparse/src/error.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/error.rs), which provides detailed variants such as **`PdfiumError`** and **`OcrError`**. When a parsing step fails, the error message itself indicates the stage because each module returns the centralized **`LiteParseError`** type. Use the variant name to decide whether to investigate PDFium bindings, OCR workers, or layout projection logic.

### Baseline Behavior with Integration Tests

The integration test in [`crates/liteparse/tests/integration_test.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/tests/integration_test.rs) runs a full parse on a known PDF and compares the JSON output. Use this file as a sanity check, or replace the fixture PDF with your problematic file and observe which assertion fails. This test is the fastest way to create a minimal reproducible case that exercises the entire pipeline from [`crates/liteparse/src/parser.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/parser.rs).

## Practical Debugging Workflow

1. **Run the CLI with maximum verbosity.** Add `--log=debug` to forward the flag to **`env_logger`**.

   ```bash
   liteparse mydoc.pdf --output json --log=debug > out.json
   ```

   The log output includes the number of raw items extracted in [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs), the rotation groups detected in `handle_rotation_reading_order`, and anchor map sizes before and after `delta_min_filter` and `intercept_filter`.

2. **Inspect the intermediate `ProjectedTextItem` vector.** In Rust, temporarily add a debug dump in [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs) after the call to **`project_items`**:

   ```rust
   let mut projected = project_items(raw_items, &config)?;
   eprintln!("PROJECTED ITEMS ({})", projected.len());
   for (i, item) in projected.iter().enumerate() {
       eprintln!("{}: {:?} – rot:{:?}", i, item.item, item.item.rotation);
   }
   ```

   This tells you whether rotation has been cleared and whether items have sensible coordinates before layout reconstruction begins.

3. **Validate OCR merging.** If you use OCR, print the OCR-only items before they reach the merge stage:

   ```rust
   let ocr_items = run_ocr(&rendered_pages)?;
   eprintln!("OCR items: {:?}", ocr_items);
   ```

   Confirm that the OCR text aligns to the same `x/y` range as the native items in [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs).

4. **Tune the flowing-text thresholds.** If whole paragraphs appear as separate lines, adjust the constants in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs).

   Try increasing `FLOWING_WIDE_LINE_RATIO` (currently `0.5`) or decreasing `FLOWING_MIN_LINE_ITEMS`. Re-run the test binary after each change and compare the output.

5. **Re-run the integration test with a custom PDF.** Swap the fixture in [`crates/liteparse/tests/integration_test.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/tests/integration_test.rs) for your problematic file and observe which assertion fails. The test prints the parsed JSON, which you can diff against the expected output to identify the exact stage where output diverges.

## Debug LiteParse from Node.js, Python, and the CLI

### Node.js Wrapper

Enable debugging and inspect the parsed structure before it reaches your application:

```ts
import { LiteParse } from "liteparse";

(async () => {
  const parser = new LiteParse({ logLevel: "debug", enableOcr: false });
  const result = await parser.parse("sample.pdf");
  console.log(JSON.stringify(result, null, 2));

  // Dump raw projected items by adding this method in packages/node/src/native.ts:
  // native.dumpProjectedItems();
})();

```

### Python Wrapper

Capture the parser output and hook into the native layer for intermediate state:

```python
from liteparse import LiteParse

parser = LiteParse(log_level="debug", ocr=False)
result = parser.parse("sample.pdf")
print(result.json(indent=2))

# To dump intermediate ProjectedTextItems, add to liteparse/parser.py:

#   from ._native import _dump_projected_items

#   _dump_projected_items()

```

### CLI Rotation Override

Force a specific rotation handling path to see how the normalizer behaves:

```bash
liteparse rotated.pdf --disable-ocr --log=debug --rotate-force=90

```

This helps you verify how `canonical_rotation` and `handle_rotation_reading_order` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) process rotated labels.

## Summary

- **Trace the pipeline:** Start at [`crates/liteparse/src/main.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/main.rs) or [`lib.rs`](https://github.com/run-llama/liteparse/blob/main/lib.rs) and follow data through [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs), [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs), and [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs).
- **Use debug logs:** Run the CLI with `--log=debug` to expose item counts, rotation groups, and anchor filter statistics.
- **Dump intermediate state:** Temporarily print the `ProjectedTextItem` vector in [`parser.rs`](https://github.com/run-llama/liteparse/blob/main/parser.rs) after `project_items` to verify coordinates and rotation.
- **Audit anchors:** Check `delta_min_filter` and `intercept_filter` in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs) when columns or tables are mis-detected.
- **Tune flowing text:** Adjust `FLOWING_WIDE_LINE_RATIO` and `FLOWING_MIN_LINE_ITEMS` at the top of [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) to fix paragraph splitting.
- **Leverage tests:** Use [`crates/liteparse/tests/integration_test.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/tests/integration_test.rs) as a reproducible baseline by swapping in your target PDF.

## Frequently Asked Questions

### How do I know which LiteParse stage is causing a parsing error?

Every stage returns a `Result<…, LiteParseError>` defined in [`crates/liteparse/src/error.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/error.rs). The error variant—such as `PdfiumError` or `OcrError`—tells you exactly which module failed, so you can inspect [`extract.rs`](https://github.com/run-llama/liteparse/blob/main/extract.rs), [`ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/ocr_merge.rs), or [`projection.rs`](https://github.com/run-llama/liteparse/blob/main/projection.rs) accordingly.

### Why are my table columns merging or splitting incorrectly in LiteParse?

Incorrect column detection usually traces back to the anchor filters in [`crates/liteparse/src/projection.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/projection.rs). `delta_min_filter` removes isolated anchors around line 549, and `intercept_filter` drops visually crossed anchors around line 995. Add debug prints before and after these filters to inspect anchor map sizes.

### How can I debug OCR-related text duplication in LiteParse?

Inspect [`crates/liteparse/src/ocr_merge.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/ocr_merge.rs) and verify that the `merged` flag is set correctly and that `item.rotation` is cleared after merging. You can also print the OCR-only items before merging to confirm their `x/y` coordinates align with the native text extracted from [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).

### What is the fastest way to reproduce a LiteParse bug for debugging?

Replace the fixture PDF in [`crates/liteparse/tests/integration_test.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/tests/integration_test.rs) with your problematic file and run the integration test. This end-to-end test compares the parsed JSON against expected output, so the first failing assertion points to the exact layout stage that diverges from expectations.