Understanding the `mcid` Field in LiteParse's TextItem: A Complete Guide

The mcid field in LiteParse's TextItem is an optional Option<i32> that stores the Marked-Content ID from a PDF's structure tree, enabling downstream applications to map extracted text back to its original logical grouping for accessibility, tagging, and selective processing.

LiteParse, developed in the run-llama/liteparse repository, exposes structural PDF metadata through the mcid field on its TextItem struct. This optional identifier captures the link between raw glyph extraction and the document's logical structure tree. For developers building accessibility tools, HTML converters, or semantic parsing pipelines, understanding the mcid field is critical to preserving document semantics beyond plain text extraction.

What the mcid Field Represents

The mcid field is defined in crates/liteparse/src/types.rs as an optional 32-bit integer:

pub mcid: Option<i32>,                     // marked‑content ID from the PDF structure tree

MCID stands for Marked-Content ID, a numeric identifier assigned by the PDF writer to elements within the PDF's structure tree. It groups together characters that belong to the same logical element, such as a paragraph, list item, or figure caption. Because not all PDFs contain marked-content elements, the field remains None when the source PDF does not provide a structure ID for a given character.

How LiteParse Extracts and Propagates mcid

LiteParse retrieves the mcid value at the PDFium layer and propagates it through the extraction pipeline into the final TextItem.

Retrieving the MCID from PDFium

In crates/pdfium/src/text_page.rs, the wrapper calls the native PDFium API FPDFPageObj_GetMarkedContentID for each page object:

let mcid = unsafe { ffi!(FPDFPageObj_GetMarkedContentID(obj)) };
if mcid >= 0 { Some(mcid) } else { None }

The wrapper then exposes this value via pdfium::TextChar::marked_content_id().

Assembling Text Segments

During segment construction in crates/liteparse/src/extract.rs, the builder captures the MCID from the first glyph of a new text segment:

self.mcid = ch.marked_content_id();

When the segment is finalized, the flush() method copies the stored value into the exported TextItem:

mcid: self.mcid,

Why the mcid Field Matters

Structural mapping allows consumers to rebuild the PDF's logical hierarchy by matching each TextItem back to its originating structure element via the MCID. Selective processing becomes possible when you filter text items by their MCID to isolate specific sections, such as text contained only within <Figure> tags. By retaining the MCID in JSON or plain-text output, downstream applications preserve the document's semantic grouping without re-parsing the original PDF.

Practical Code Examples

Accessing mcid in Rust

After parsing a document, iterate over pages and text items to read the optional mcid field:

use liteparse::parser::LiteParse;

let parser = LiteParse::new("sample.pdf");
let parsed = parser.parse().await?;
for page in parsed.pages {
    for item in page.text_items {
        if let Some(mcid) = item.mcid {
            println!("MCID {} → '{}'", mcid, item.text);
        }
    }
}

Reading mcid from Python JSON Output

When using LiteParse's Python bindings, the MCID appears as an integer key in the serialized text item dictionary:

from liteparse import LiteParse

parser = LiteParse("sample.pdf")
result = parser.parse()
for page in result["pages"]:
    for ti in page["text_items"]:
        if "mcid" in ti:
            print(f"MCID {ti['mcid']}: {ti['text']}")

JSON Output Structure

Each TextItem exported to JSON includes the mcid key only when the value is present:

{
  "text": "Hello world",
  "x": 12.3,
  "y": 45.6,
  "width": 50.0,
  "height": 10.0,
  "mcid": 7,
  "font_name": "Helvetica",
  "font_size": 12.0
}

Summary

  • The mcid field on TextItem is an Option<i32> representing the PDF's Marked-Content ID from the structure tree.
  • LiteParse extracts the MCID through PDFium's FPDFPageObj_GetMarkedContentID API in crates/pdfium/src/text_page.rs.
  • The builder propagates the MCID from TextChar into TextItem during segment assembly in crates/liteparse/src/extract.rs.
  • Downstream workflows use mcid to reconstruct logical document structure, filter content by semantic group, and preserve accessibility metadata.

Frequently Asked Questions

What does the mcid field represent in LiteParse?

The mcid field represents the Marked-Content ID, an optional numeric identifier that links extracted text to its originating element in the PDF's logical structure tree. It enables semantic grouping of characters that belong to the same structural element, such as a paragraph or caption.

Is the mcid field always present on every TextItem?

No, the mcid field is optional and typed as Option<i32>. If the source PDF does not define marked-content IDs for a given text region, LiteParse sets the field to None, and the key is omitted from JSON output.

How does LiteParse retrieve the mcid from a PDF?

LiteParse calls PDFium's native FPDFPageObj_GetMarkedContentID function inside crates/pdfium/src/text_page.rs. The resulting value is propagated through TextChar::marked_content_id() and stored during segment building in crates/liteparse/src/extract.rs.

How can developers use the mcid field in downstream applications?

Developers can use mcid to map extracted text back to logical PDF structure elements for accessibility tools, filter TextItem collections by specific structural sections, and retain semantic grouping information in custom JSON or HTML conversion pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →