Debugging LiteParse Parsing Issues: A Step-by-Step Source Code Guide
You can debug LiteParse parsing issues by tracing data through eight well-defined pipeline stages—from raw PDF extraction in extract.rs to layout reconstruction in projection.rs—using CLI debug logs, intermediate ProjectedTextItem dumps, and the integration test suite as a reproducible baseline.
Follow the flow of data through run-llama/liteparse to pinpoint exactly where layout reconstruction diverges from expectations. The crate processes documents through a strict pipeline, and each stage exposes distinct inspection points in the source code. Understanding these stages is the fastest way to resolve misaligned columns, missing paragraphs, or incorrect rotations.
Understand the LiteParse Pipeline
LiteParse processes PDFs through a strict sequence of discrete transformations. The journey starts at the CLI entry point in crates/liteparse/src/main.rs and the public LiteParse struct in crates/liteparse/src/lib.rs, then moves through raw extraction, rotation normalization, anchor mapping, flowing-text classification, and optional OCR merging. Every stage returns a Result type defined in crates/liteparse/src/error.rs, so you can isolate failures by examining the module where the error originates.
Debug LiteParse Parsing Issues by Pipeline Stage
Verify the CLI or API Entry Point
Start by confirming that your LiteParseConfig is initialized correctly. In crates/liteparse/src/main.rs, the CLI creates the configuration and hands it to parser::parse, while crates/liteparse/src/lib.rs exposes the same public API for programmatic use. If the output is unexpectedly empty, ensure that your PDF path and flags are reaching the internal parser.
Inspect Raw Text Extraction
Before layout reconstruction begins, PDFium returns a flat list of TextItem objects in crates/liteparse/src/extract.rs. Verify that the raw list contains the expected number of items, correct coordinates, and text strings by checking the extract_page implementation. If the raw extraction is missing text, the problem is upstream of layout logic and likely stems from the PDFium binding or the input file itself.
Check Rotation Normalization
Rotated text is normalized before layout reconstruction in crates/liteparse/src/projection.rs. Look at canonical_rotation, which rounds angles to 0/90/180/270°, and handle_rotation_reading_order, which groups items by rotation and re-orders them. When text appears out of order on rotated pages, inspect these functions around lines 89 and 113 to confirm that rotation values are being bucketed and transformed correctly.
Audit Anchor Detection and Filtering
The engine builds left, right, and center anchor maps to snap boxes together, defined in crates/liteparse/src/projection.rs. Two filters are common culprits when columns are mis-detected: delta_min_filter (removes isolated anchors) around line 549, and intercept_filter (drops anchors that are visually crossed) around line 995. If your table columns are merging or splitting incorrectly, add debug prints before and after these filters to inspect anchor map sizes.
Tune Flowing-Text Constants
A set of FLOWING_* constants at the top of crates/liteparse/src/projection.rs controls when a block is treated as flowing prose rather than a table. FLOWING_MAX_TOTAL_ANCHORS and FLOWING_WIDE_LINE_RATIO (currently 0.5) determine paragraph boundaries. If the parser mistakenly splits a paragraph into separate lines, try increasing FLOWING_WIDE_LINE_RATIO or decreasing FLOWING_MIN_LINE_ITEMS in a local copy and re-running the test binary.
Validate OCR Result Merging
When OCR is enabled, crates/liteparse/src/ocr_merge.rs combines native PDF text with OCR-generated text. Verify that the merged flag is set correctly and that item.rotation is cleared after merging. If you see duplicated or misaligned text, print the OCR-only items before merging to confirm they align to the same x/y range as the native items.
Read Error Variants for Stage Identification
All stages propagate failures through crates/liteparse/src/error.rs, which provides detailed variants such as PdfiumError and OcrError. When a parsing step fails, the error message itself indicates the stage because each module returns the centralized LiteParseError type. Use the variant name to decide whether to investigate PDFium bindings, OCR workers, or layout projection logic.
Baseline Behavior with Integration Tests
The integration test in crates/liteparse/tests/integration_test.rs runs a full parse on a known PDF and compares the JSON output. Use this file as a sanity check, or replace the fixture PDF with your problematic file and observe which assertion fails. This test is the fastest way to create a minimal reproducible case that exercises the entire pipeline from crates/liteparse/src/parser.rs.
Practical Debugging Workflow
-
Run the CLI with maximum verbosity. Add
--log=debugto forward the flag toenv_logger.liteparse mydoc.pdf --output json --log=debug > out.jsonThe log output includes the number of raw items extracted in
extract.rs, the rotation groups detected inhandle_rotation_reading_order, and anchor map sizes before and afterdelta_min_filterandintercept_filter. -
Inspect the intermediate
ProjectedTextItemvector. In Rust, temporarily add a debug dump inparser.rsafter the call toproject_items:let mut projected = project_items(raw_items, &config)?; eprintln!("PROJECTED ITEMS ({})", projected.len()); for (i, item) in projected.iter().enumerate() { eprintln!("{}: {:?} – rot:{:?}", i, item.item, item.item.rotation); }This tells you whether rotation has been cleared and whether items have sensible coordinates before layout reconstruction begins.
-
Validate OCR merging. If you use OCR, print the OCR-only items before they reach the merge stage:
let ocr_items = run_ocr(&rendered_pages)?; eprintln!("OCR items: {:?}", ocr_items);Confirm that the OCR text aligns to the same
x/yrange as the native items incrates/liteparse/src/ocr_merge.rs. -
Tune the flowing-text thresholds. If whole paragraphs appear as separate lines, adjust the constants in
crates/liteparse/src/projection.rs.Try increasing
FLOWING_WIDE_LINE_RATIO(currently0.5) or decreasingFLOWING_MIN_LINE_ITEMS. Re-run the test binary after each change and compare the output. -
Re-run the integration test with a custom PDF. Swap the fixture in
crates/liteparse/tests/integration_test.rsfor your problematic file and observe which assertion fails. The test prints the parsed JSON, which you can diff against the expected output to identify the exact stage where output diverges.
Debug LiteParse from Node.js, Python, and the CLI
Node.js Wrapper
Enable debugging and inspect the parsed structure before it reaches your application:
import { LiteParse } from "liteparse";
(async () => {
const parser = new LiteParse({ logLevel: "debug", enableOcr: false });
const result = await parser.parse("sample.pdf");
console.log(JSON.stringify(result, null, 2));
// Dump raw projected items by adding this method in packages/node/src/native.ts:
// native.dumpProjectedItems();
})();
Python Wrapper
Capture the parser output and hook into the native layer for intermediate state:
from liteparse import LiteParse
parser = LiteParse(log_level="debug", ocr=False)
result = parser.parse("sample.pdf")
print(result.json(indent=2))
# To dump intermediate ProjectedTextItems, add to liteparse/parser.py:
# from ._native import _dump_projected_items
# _dump_projected_items()
CLI Rotation Override
Force a specific rotation handling path to see how the normalizer behaves:
liteparse rotated.pdf --disable-ocr --log=debug --rotate-force=90
This helps you verify how canonical_rotation and handle_rotation_reading_order in crates/liteparse/src/projection.rs process rotated labels.
Summary
- Trace the pipeline: Start at
crates/liteparse/src/main.rsorlib.rsand follow data throughextract.rs,projection.rs, andocr_merge.rs. - Use debug logs: Run the CLI with
--log=debugto expose item counts, rotation groups, and anchor filter statistics. - Dump intermediate state: Temporarily print the
ProjectedTextItemvector inparser.rsafterproject_itemsto verify coordinates and rotation. - Audit anchors: Check
delta_min_filterandintercept_filterincrates/liteparse/src/projection.rswhen columns or tables are mis-detected. - Tune flowing text: Adjust
FLOWING_WIDE_LINE_RATIOandFLOWING_MIN_LINE_ITEMSat the top ofprojection.rsto fix paragraph splitting. - Leverage tests: Use
crates/liteparse/tests/integration_test.rsas a reproducible baseline by swapping in your target PDF.
Frequently Asked Questions
How do I know which LiteParse stage is causing a parsing error?
Every stage returns a Result<…, LiteParseError> defined in crates/liteparse/src/error.rs. The error variant—such as PdfiumError or OcrError—tells you exactly which module failed, so you can inspect extract.rs, ocr_merge.rs, or projection.rs accordingly.
Why are my table columns merging or splitting incorrectly in LiteParse?
Incorrect column detection usually traces back to the anchor filters in crates/liteparse/src/projection.rs. delta_min_filter removes isolated anchors around line 549, and intercept_filter drops visually crossed anchors around line 995. Add debug prints before and after these filters to inspect anchor map sizes.
How can I debug OCR-related text duplication in LiteParse?
Inspect crates/liteparse/src/ocr_merge.rs and verify that the merged flag is set correctly and that item.rotation is cleared after merging. You can also print the OCR-only items before merging to confirm their x/y coordinates align with the native text extracted from crates/liteparse/src/extract.rs.
What is the fastest way to reproduce a LiteParse bug for debugging?
Replace the fixture PDF in crates/liteparse/tests/integration_test.rs with your problematic file and run the integration test. This end-to-end test compares the parsed JSON against expected output, so the first failing assertion points to the exact layout stage that diverges from expectations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →