How to Perform Document Ingestion with Docling and Marker in Sieves

Use the Docling or Marker tasks in Sieves to convert PDFs and documents into plain text by passing Doc objects with file URIs, which automatically extract content into doc.text for downstream processing.

Sieves provides dedicated pre-processing ingestion tasks that wrap Docling and Marker, two popular open-source document parsing libraries. These tasks handle the heavy lifting of converting binary documents into structured text within the sieves framework, enabling seamless integration with chunking, embedding, and predictive modeling pipelines.

Understanding Document Ingestion in Sieves

Both Docling and Marker inherit from sieves.tasks.core.Task and follow a standardized five-step execution model defined in sieves/tasks/core.py:

  1. Receive an iterable of sieves.data.Doc objects, each requiring a .uri attribute pointing to the source file.
  2. Validate incoming documents to ensure URIs are present and warn if existing .text will be overwritten.
  3. Delegate parsing to the external library (docling.document_converter.DocumentConverter or marker.converters.pdf.PdfConverter).
  4. Store extracted content in doc.text and optionally attach metadata or images.
  5. Yield enriched Doc objects for downstream consumption.

Docling vs. Marker: Key Differences

Feature Docling Ingestion Marker Ingestion
Core Class Docling(Task) in sieves/tasks/preprocessing/ingestion/docling_.py Marker(Task) in sieves/tasks/preprocessing/ingestion/marker_.py
Supported Formats PDF, DOCX, and other formats via DocumentConverter PDF-focused via PdfConverter/TableConverter
Export Formats markdown (default), html markdown (default), html, json
Batch Handling Auto-adjusts docling.datamodel.settings.settings.perf.doc_batch_size based on task batch_size No internal batch tuning; relies on task iteration
Image Extraction Available via Docling's native capabilities Native support via extract_images=True parameter
Metadata Storage Stores raw ParsedResource in doc.meta[task_id] when include_meta=True Stores Converter instance in serialized state; images attach to doc.images

Installation and Setup

Both tasks are available as optional extras. Install them using the ingestion extra or directly install the underlying libraries:


# Install via sieves extras

pip install sieves[ingestion]

# Or install parsers directly

pip install docling
pip install marker

Import the tasks from their respective modules:

from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.ingestion.marker_ import Marker
from sieves.data import Doc

Performing Document Ingestion with Docling

Docling converts documents into markdown or HTML format using IBM's DocumentConverter. The task automatically handles batch processing and integrates with Docling's performance settings.

Basic Docling Ingestion

Create a Doc with a file URI and process it through the Docling task:

from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling

# Create a Doc pointing at a local file

doc = Doc(uri="file:///path/to/report.pdf")

# Initialize the Docling task with markdown output (default)

docling_task = Docling(export_format="markdown")

# Run the task

processed_docs = list(docling_task([doc]))

# Access extracted text

print(processed_docs[0].text)  # Markdown string

# Access metadata if include_meta=True

# print(processed_docs[0].meta)  # {'docling': ParsedResource(...)}

The constructor and _call implementation are defined in lines 23-84 of sieves/tasks/preprocessing/ingestion/docling_.py.

Advanced Docling with Batching and Filtering

Use the batch_size parameter to control throughput and condition to filter documents before processing:

from pathlib import Path
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling

# Build a collection of Docs from a directory

docs = [Doc(uri=str(p)) for p in Path("data/reports").glob("*.pdf")]

# Filter: only process PDFs larger than 1 MiB

def large_file(doc: Doc) -> bool:
    return Path(doc.uri).stat().st_size > 1_048_576

# Configure task with HTML output and custom batching

docling = Docling(
    export_format="html",
    batch_size=10,           # Process in batches of 10

    condition=large_file,      # Skip tiny files

    include_meta=True,
)

processed = list(docling(docs))
print(f"Parsed {len(processed)} large documents.")

The batch size is applied in __init__ (lines 44-48) and the condition argument is passed to the base Task (line 40) in sieves/tasks/preprocessing/ingestion/docling_.py.

Performing Document Ingestion with Marker

Marker specializes in PDF-to-markdown conversion with optional image extraction. It supports multiple output formats including markdown, HTML, and JSON.

Marker with Image Extraction

Process PDFs and extract embedded images into the Doc.images attribute:

from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.marker_ import Marker

doc = Doc(uri="file:///path/to/paper.pdf")

marker = Marker(
    export_format="markdown",
    extract_images=True,   # Store extracted images in doc.images

    include_meta=False,
)

# Marker yields an iterator; take the first result

doc = next(marker([doc]))

print(doc.text)            # Markdown representation of the PDF

print(f"Found {len(doc.images)} images.")  # Images extracted by Marker

Image extraction is performed in the _call method at lines 9-12 of the loop in sieves/tasks/preprocessing/ingestion/marker_.py.

Integrating Ingestion into a Pipeline

Ingestion tasks work seamlessly within sieves.Pipeline objects, allowing you to chain document parsing with chunking and predictive modeling:

from sieves import Pipeline
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.chunking.chonkie_ import Chonkie
import chonkie, tokenizers

# Build pipeline: Docling → Token-based chunking

pipeline = Pipeline([
    Docling(export_format="markdown"),
    Chonkie(chunker=chonkie.TokenChunker(tokenizers.Tokenizer.from_pretrained("gpt2")))
])

# Run on multiple documents

docs = [Doc(uri="file:///data/file1.pdf"), Doc(uri="file:///data/file2.pdf")]
processed = pipeline(docs)

for d in processed:
    print(f"{d.uri} → {len(d.chunks)} chunks")

The Pipeline automatically respects each task's batch_size and caching semantics defined in sieves/tasks/core.py.

Summary

  • Docling and Marker tasks in Sieves wrap third-party parsers to convert binary documents into structured text within the Doc.text attribute.
  • Both tasks inherit from sieves.tasks.core.Task and follow the standard execution model: validate, delegate, store, yield.
  • Docling (sieves/tasks/preprocessing/ingestion/docling_.py) supports multiple document formats (PDF, DOCX) with markdown/HTML export and automatic batch size tuning.
  • Marker (sieves/tasks/preprocessing/ingestion/marker_.py) specializes in PDF conversion with optional image extraction into doc.images and supports markdown/HTML/JSON output formats.
  • Install via pip install sieves[ingestion] or directly install docling/marker.
  • Both tasks integrate seamlessly into sieves.Pipeline objects for end-to-end document processing workflows.

Frequently Asked Questions

What file formats are supported by Docling and Marker in Sieves?

Docling supports PDF, DOCX, and other office formats through its DocumentConverter class, converting them to markdown or HTML. Marker focuses specifically on PDF files using PdfConverter or TableConverter, with output options including markdown, HTML, and JSON. Both tasks require documents to be provided as Doc objects with valid file URIs in the .uri attribute.

How do I extract images from PDFs using the Marker task?

Set the extract_images=True parameter when initializing the Marker task. When processing documents, Marker will extract embedded images and store them in the doc.images list of each processed Doc object. This occurs in the _call method of sieves/tasks/preprocessing/ingestion/marker_.py. Note that image extraction is specific to Marker; Docling handles images differently through its own document object model.

Can I filter which documents get processed by the ingestion tasks?

Yes, both Docling and Marker accept a condition callable parameter inherited from the base Task class in sieves/tasks/core.py. This function receives a Doc object and returns a boolean indicating whether to process it. For example, you can filter by file size, extension, or URI pattern. The task skips any documents where the condition returns False, allowing you to process only relevant files in large document collections.

What is the difference between the batch_size parameter in Docling versus Marker?

In Docling, the batch_size parameter automatically adjusts the underlying library's performance settings by modifying docling.datamodel.settings.settings.perf.doc_batch_size to optimize throughput for the specified batch size. In Marker, the batch_size parameter is handled by the base Task class and simply controls how many documents are passed to the task's _call method at once, without modifying internal Marker settings. Both respect the standard Sieves task execution model defined in sieves/tasks/core.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →