How to Enable or Disable Link Extraction in LiteParse Markdown Output

To disable link extraction in LiteParse markdown output, pass the --no-links CLI flag or set extract_links to false in LiteParseConfig.

LiteParse, an open-source PDF parsing library developed by run-llama, converts PDF documents into structured formats including Markdown. By default, it extracts URI hyperlink annotations and renders them as inline Markdown links ([text](url)). Controlling this behavior is essential when you need clean text without embedded URLs or want to preserve clickable references in your output.

Link extraction is handled by the assign_links routine in crates/liteparse/src/extract.rs. This function attaches URI annotations to TextItem objects during the parsing phase. When enabled, LiteParse wraps the detected text and its corresponding URL into standard Markdown link syntax. This transformation only occurs when the output format is set to Markdown; other formats such as JSON or plain text ignore the link extraction setting entirely.

You can toggle link extraction using either the command-line interface or the Rust API. Both methods rely on the extract_links boolean field defined in LiteParseConfig.

Command-Line Interface

The CLI provides a --no-links flag that flips the extract_links configuration to false. By default, the flag is absent and extraction is enabled.


# Disable link extraction (default is enabled)

liteparse parse mydoc.pdf --output-format markdown --no-links

# Explicitly enable link extraction (redundant, but explicit)

liteparse parse mydoc.pdf --output-format markdown

The flag is defined in crates/liteparse/src/main.rs within the no_links field (lines 16–22).

Rust API

When constructing a LiteParseConfig programmatically, set the extract_links field directly before instantiating the parser.

use liteparse::config::{LiteParseConfig, OutputFormat};

// Disable link extraction
let mut cfg = LiteParseConfig::default();
cfg.extract_links = false;
cfg.output_format = OutputFormat::Markdown;

let parser = liteparse::LiteParse::new(cfg);
let result = parser.parse_input(liteparse::PdfInput::Path("mydoc.pdf".into())).await?;
println!("{}", result.markdown); // Outputs Markdown without [text](url)
// Explicitly enable link extraction (default behavior)
let mut cfg = LiteParseConfig::default();
cfg.extract_links = true;
cfg.output_format = OutputFormat::Markdown;

The default value for extract_links is true, as defined in crates/liteparse/src/config.rs (lines 38–40).

Implementation Details and Conditional Logic

Link extraction is conditionally applied inside LiteParse::parse_input in crates/liteparse/src/parser.rs (lines 2008–2011). The parser checks both the output format and the extract_links flag before invoking the assign_links routine. Consequently, setting extract_links to false prevents the parser from emitting Markdown links even if the PDF contains valid URI annotations.

Key source files:

Summary

  • Default behavior: Link extraction is enabled (extract_links = true), converting PDF annotations into Markdown links.
  • CLI toggle: Use --no-links to disable extraction when running liteparse parse.
  • API toggle: Set cfg.extract_links = false in LiteParseConfig before creating the parser instance.
  • Format restriction: Link extraction only affects Markdown output; JSON and plain text ignore the setting.
  • Source authority: Configuration resides in config.rs, CLI logic in main.rs, and conditional parsing in parser.rs.

Frequently Asked Questions

No. The --no-links flag and the underlying extract_links configuration only apply when the output format is Markdown. In crates/liteparse/src/parser.rs, the link extraction guard checks the output format before processing annotations, so JSON and plain text outputs remain unchanged regardless of this setting.

No. LiteParse applies the extract_links setting globally to the entire document during the parsing phase. There is no per-page granularity in the current implementation; the flag is evaluated once in LiteParse::parse_input and applies to all extracted text items uniformly.

What happens if a PDF contains malformed or invalid URIs?

The assign_links routine in crates/liteparse/src/extract.rs attaches URI annotations as-is. LiteParse does not validate URL syntax during extraction; it simply maps the annotation text to the associated URI string. If the PDF contains malformed URIs, they will appear verbatim in the Markdown output (e.g., [text](invalid-url)).

No. As of the current codebase, extract_links is only configurable via the --no-links CLI flag or programmatically through the LiteParseConfig struct. There is no built-in support for environment variable or file-based configuration for this specific option in the analyzed source files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →