# Understanding the `mcid` Field in LiteParse's TextItem: A Complete Guide

> Discover the mcid field in LiteParse TextItem. Learn how this optional field maps text to PDF logical structures for accessibility and processing.

- Repository: [LlamaIndex/liteparse](https://github.com/run-llama/liteparse)
- Tags: deep-dive
- Published: 2026-06-06

---

**The `mcid` field in LiteParse's `TextItem` is an optional `Option<i32>` that stores the Marked-Content ID from a PDF's structure tree, enabling downstream applications to map extracted text back to its original logical grouping for accessibility, tagging, and selective processing.**

LiteParse, developed in the `run-llama/liteparse` repository, exposes structural PDF metadata through the `mcid` field on its `TextItem` struct. This optional identifier captures the link between raw glyph extraction and the document's logical structure tree. For developers building accessibility tools, HTML converters, or semantic parsing pipelines, understanding the `mcid` field is critical to preserving document semantics beyond plain text extraction.

## What the `mcid` Field Represents

The `mcid` field is defined in [`crates/liteparse/src/types.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/types.rs) as an optional 32-bit integer:

```rust
pub mcid: Option<i32>,                     // marked‑content ID from the PDF structure tree

```

**MCID** stands for **Marked-Content ID**, a numeric identifier assigned by the PDF writer to elements within the PDF's structure tree. It groups together characters that belong to the same logical element, such as a paragraph, list item, or figure caption. Because not all PDFs contain marked-content elements, the field remains `None` when the source PDF does not provide a structure ID for a given character.

## How LiteParse Extracts and Propagates `mcid`

LiteParse retrieves the `mcid` value at the PDFium layer and propagates it through the extraction pipeline into the final `TextItem`.

### Retrieving the MCID from PDFium

In [`crates/pdfium/src/text_page.rs`](https://github.com/run-llama/liteparse/blob/main/crates/pdfium/src/text_page.rs), the wrapper calls the native PDFium API `FPDFPageObj_GetMarkedContentID` for each page object:

```rust
let mcid = unsafe { ffi!(FPDFPageObj_GetMarkedContentID(obj)) };
if mcid >= 0 { Some(mcid) } else { None }

```

The wrapper then exposes this value via `pdfium::TextChar::marked_content_id()`.

### Assembling Text Segments

During segment construction in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs), the builder captures the MCID from the first glyph of a new text segment:

```rust
self.mcid = ch.marked_content_id();

```

When the segment is finalized, the `flush()` method copies the stored value into the exported `TextItem`:

```rust
mcid: self.mcid,

```

## Why the `mcid` Field Matters

Structural mapping allows consumers to rebuild the PDF's logical hierarchy by matching each `TextItem` back to its originating structure element via the MCID. **Selective processing** becomes possible when you filter text items by their MCID to isolate specific sections, such as text contained only within `<Figure>` tags. By retaining the MCID in JSON or plain-text output, downstream applications preserve the document's semantic grouping without re-parsing the original PDF.

## Practical Code Examples

### Accessing `mcid` in Rust

After parsing a document, iterate over pages and text items to read the optional `mcid` field:

```rust
use liteparse::parser::LiteParse;

let parser = LiteParse::new("sample.pdf");
let parsed = parser.parse().await?;
for page in parsed.pages {
    for item in page.text_items {
        if let Some(mcid) = item.mcid {
            println!("MCID {} → '{}'", mcid, item.text);
        }
    }
}

```

### Reading `mcid` from Python JSON Output

When using LiteParse's Python bindings, the MCID appears as an integer key in the serialized text item dictionary:

```python
from liteparse import LiteParse

parser = LiteParse("sample.pdf")
result = parser.parse()
for page in result["pages"]:
    for ti in page["text_items"]:
        if "mcid" in ti:
            print(f"MCID {ti['mcid']}: {ti['text']}")

```

### JSON Output Structure

Each `TextItem` exported to JSON includes the `mcid` key only when the value is present:

```json
{
  "text": "Hello world",
  "x": 12.3,
  "y": 45.6,
  "width": 50.0,
  "height": 10.0,
  "mcid": 7,
  "font_name": "Helvetica",
  "font_size": 12.0
}

```

## Summary

- The `mcid` field on `TextItem` is an `Option<i32>` representing the PDF's Marked-Content ID from the structure tree.
- LiteParse extracts the MCID through PDFium's `FPDFPageObj_GetMarkedContentID` API in [`crates/pdfium/src/text_page.rs`](https://github.com/run-llama/liteparse/blob/main/crates/pdfium/src/text_page.rs).
- The builder propagates the MCID from `TextChar` into `TextItem` during segment assembly in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).
- Downstream workflows use `mcid` to reconstruct logical document structure, filter content by semantic group, and preserve accessibility metadata.

## Frequently Asked Questions

### What does the `mcid` field represent in LiteParse?

The `mcid` field represents the **Marked-Content ID**, an optional numeric identifier that links extracted text to its originating element in the PDF's logical structure tree. It enables semantic grouping of characters that belong to the same structural element, such as a paragraph or caption.

### Is the `mcid` field always present on every `TextItem`?

No, the `mcid` field is optional and typed as `Option<i32>`. If the source PDF does not define marked-content IDs for a given text region, LiteParse sets the field to `None`, and the key is omitted from JSON output.

### How does LiteParse retrieve the `mcid` from a PDF?

LiteParse calls PDFium's native `FPDFPageObj_GetMarkedContentID` function inside [`crates/pdfium/src/text_page.rs`](https://github.com/run-llama/liteparse/blob/main/crates/pdfium/src/text_page.rs). The resulting value is propagated through `TextChar::marked_content_id()` and stored during segment building in [`crates/liteparse/src/extract.rs`](https://github.com/run-llama/liteparse/blob/main/crates/liteparse/src/extract.rs).

### How can developers use the `mcid` field in downstream applications?

Developers can use `mcid` to map extracted text back to logical PDF structure elements for accessibility tools, filter `TextItem` collections by specific structural sections, and retain semantic grouping information in custom JSON or HTML conversion pipelines.