How to Enable or Disable Link Extraction in LiteParse Markdown Output
To disable link extraction in LiteParse markdown output, pass the --no-links CLI flag or set extract_links to false in LiteParseConfig.
LiteParse, an open-source PDF parsing library developed by run-llama, converts PDF documents into structured formats including Markdown. By default, it extracts URI hyperlink annotations and renders them as inline Markdown links ([text](url)). Controlling this behavior is essential when you need clean text without embedded URLs or want to preserve clickable references in your output.
Understanding Link Extraction in LiteParse
Link extraction is handled by the assign_links routine in crates/liteparse/src/extract.rs. This function attaches URI annotations to TextItem objects during the parsing phase. When enabled, LiteParse wraps the detected text and its corresponding URL into standard Markdown link syntax. This transformation only occurs when the output format is set to Markdown; other formats such as JSON or plain text ignore the link extraction setting entirely.
Configuration Methods to Control Link Extraction
You can toggle link extraction using either the command-line interface or the Rust API. Both methods rely on the extract_links boolean field defined in LiteParseConfig.
Command-Line Interface
The CLI provides a --no-links flag that flips the extract_links configuration to false. By default, the flag is absent and extraction is enabled.
# Disable link extraction (default is enabled)
liteparse parse mydoc.pdf --output-format markdown --no-links
# Explicitly enable link extraction (redundant, but explicit)
liteparse parse mydoc.pdf --output-format markdown
The flag is defined in crates/liteparse/src/main.rs within the no_links field (lines 16–22).
Rust API
When constructing a LiteParseConfig programmatically, set the extract_links field directly before instantiating the parser.
use liteparse::config::{LiteParseConfig, OutputFormat};
// Disable link extraction
let mut cfg = LiteParseConfig::default();
cfg.extract_links = false;
cfg.output_format = OutputFormat::Markdown;
let parser = liteparse::LiteParse::new(cfg);
let result = parser.parse_input(liteparse::PdfInput::Path("mydoc.pdf".into())).await?;
println!("{}", result.markdown); // Outputs Markdown without [text](url)
// Explicitly enable link extraction (default behavior)
let mut cfg = LiteParseConfig::default();
cfg.extract_links = true;
cfg.output_format = OutputFormat::Markdown;
The default value for extract_links is true, as defined in crates/liteparse/src/config.rs (lines 38–40).
Implementation Details and Conditional Logic
Link extraction is conditionally applied inside LiteParse::parse_input in crates/liteparse/src/parser.rs (lines 2008–2011). The parser checks both the output format and the extract_links flag before invoking the assign_links routine. Consequently, setting extract_links to false prevents the parser from emitting Markdown links even if the PDF contains valid URI annotations.
Key source files:
crates/liteparse/src/config.rs— DefinesLiteParseConfig::extract_linkswith a default oftrue.crates/liteparse/src/main.rs— Implements the--no-linksCLI flag.crates/liteparse/src/parser.rs— Guards link extraction based on format and configuration.crates/liteparse/src/extract.rs— Contains theassign_linksfunction that attaches URIs to text items.
Summary
- Default behavior: Link extraction is enabled (
extract_links = true), converting PDF annotations into Markdown links. - CLI toggle: Use
--no-linksto disable extraction when runningliteparse parse. - API toggle: Set
cfg.extract_links = falseinLiteParseConfigbefore creating the parser instance. - Format restriction: Link extraction only affects Markdown output; JSON and plain text ignore the setting.
- Source authority: Configuration resides in
config.rs, CLI logic inmain.rs, and conditional parsing inparser.rs.
Frequently Asked Questions
Does the --no-links flag affect JSON or plain text output?
No. The --no-links flag and the underlying extract_links configuration only apply when the output format is Markdown. In crates/liteparse/src/parser.rs, the link extraction guard checks the output format before processing annotations, so JSON and plain text outputs remain unchanged regardless of this setting.
Can I enable link extraction for a single PDF page while disabling it for others?
No. LiteParse applies the extract_links setting globally to the entire document during the parsing phase. There is no per-page granularity in the current implementation; the flag is evaluated once in LiteParse::parse_input and applies to all extracted text items uniformly.
What happens if a PDF contains malformed or invalid URIs?
The assign_links routine in crates/liteparse/src/extract.rs attaches URI annotations as-is. LiteParse does not validate URL syntax during extraction; it simply maps the annotation text to the associated URI string. If the PDF contains malformed URIs, they will appear verbatim in the Markdown output (e.g., [text](invalid-url)).
Is extract_links available in environment variables or configuration files?
No. As of the current codebase, extract_links is only configurable via the --no-links CLI flag or programmatically through the LiteParseConfig struct. There is no built-in support for environment variable or file-based configuration for this specific option in the analyzed source files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →