How to Perform Document Ingestion with Docling and Marker in Sieves
Use the Docling or Marker tasks in Sieves to convert PDFs and documents into plain text by passing Doc objects with file URIs, which automatically extract content into doc.text for downstream processing.
Sieves provides dedicated pre-processing ingestion tasks that wrap Docling and Marker, two popular open-source document parsing libraries. These tasks handle the heavy lifting of converting binary documents into structured text within the sieves framework, enabling seamless integration with chunking, embedding, and predictive modeling pipelines.
Understanding Document Ingestion in Sieves
Both Docling and Marker inherit from sieves.tasks.core.Task and follow a standardized five-step execution model defined in sieves/tasks/core.py:
- Receive an iterable of
sieves.data.Docobjects, each requiring a.uriattribute pointing to the source file. - Validate incoming documents to ensure URIs are present and warn if existing
.textwill be overwritten. - Delegate parsing to the external library (
docling.document_converter.DocumentConverterormarker.converters.pdf.PdfConverter). - Store extracted content in
doc.textand optionally attach metadata or images. - Yield enriched
Docobjects for downstream consumption.
Docling vs. Marker: Key Differences
| Feature | Docling Ingestion | Marker Ingestion |
|---|---|---|
| Core Class | Docling(Task) in sieves/tasks/preprocessing/ingestion/docling_.py |
Marker(Task) in sieves/tasks/preprocessing/ingestion/marker_.py |
| Supported Formats | PDF, DOCX, and other formats via DocumentConverter |
PDF-focused via PdfConverter/TableConverter |
| Export Formats | markdown (default), html |
markdown (default), html, json |
| Batch Handling | Auto-adjusts docling.datamodel.settings.settings.perf.doc_batch_size based on task batch_size |
No internal batch tuning; relies on task iteration |
| Image Extraction | Available via Docling's native capabilities | Native support via extract_images=True parameter |
| Metadata Storage | Stores raw ParsedResource in doc.meta[task_id] when include_meta=True |
Stores Converter instance in serialized state; images attach to doc.images |
Installation and Setup
Both tasks are available as optional extras. Install them using the ingestion extra or directly install the underlying libraries:
# Install via sieves extras
pip install sieves[ingestion]
# Or install parsers directly
pip install docling
pip install marker
Import the tasks from their respective modules:
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.ingestion.marker_ import Marker
from sieves.data import Doc
Performing Document Ingestion with Docling
Docling converts documents into markdown or HTML format using IBM's DocumentConverter. The task automatically handles batch processing and integrates with Docling's performance settings.
Basic Docling Ingestion
Create a Doc with a file URI and process it through the Docling task:
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
# Create a Doc pointing at a local file
doc = Doc(uri="file:///path/to/report.pdf")
# Initialize the Docling task with markdown output (default)
docling_task = Docling(export_format="markdown")
# Run the task
processed_docs = list(docling_task([doc]))
# Access extracted text
print(processed_docs[0].text) # Markdown string
# Access metadata if include_meta=True
# print(processed_docs[0].meta) # {'docling': ParsedResource(...)}
The constructor and _call implementation are defined in lines 23-84 of sieves/tasks/preprocessing/ingestion/docling_.py.
Advanced Docling with Batching and Filtering
Use the batch_size parameter to control throughput and condition to filter documents before processing:
from pathlib import Path
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
# Build a collection of Docs from a directory
docs = [Doc(uri=str(p)) for p in Path("data/reports").glob("*.pdf")]
# Filter: only process PDFs larger than 1 MiB
def large_file(doc: Doc) -> bool:
return Path(doc.uri).stat().st_size > 1_048_576
# Configure task with HTML output and custom batching
docling = Docling(
export_format="html",
batch_size=10, # Process in batches of 10
condition=large_file, # Skip tiny files
include_meta=True,
)
processed = list(docling(docs))
print(f"Parsed {len(processed)} large documents.")
The batch size is applied in __init__ (lines 44-48) and the condition argument is passed to the base Task (line 40) in sieves/tasks/preprocessing/ingestion/docling_.py.
Performing Document Ingestion with Marker
Marker specializes in PDF-to-markdown conversion with optional image extraction. It supports multiple output formats including markdown, HTML, and JSON.
Marker with Image Extraction
Process PDFs and extract embedded images into the Doc.images attribute:
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.marker_ import Marker
doc = Doc(uri="file:///path/to/paper.pdf")
marker = Marker(
export_format="markdown",
extract_images=True, # Store extracted images in doc.images
include_meta=False,
)
# Marker yields an iterator; take the first result
doc = next(marker([doc]))
print(doc.text) # Markdown representation of the PDF
print(f"Found {len(doc.images)} images.") # Images extracted by Marker
Image extraction is performed in the _call method at lines 9-12 of the loop in sieves/tasks/preprocessing/ingestion/marker_.py.
Integrating Ingestion into a Pipeline
Ingestion tasks work seamlessly within sieves.Pipeline objects, allowing you to chain document parsing with chunking and predictive modeling:
from sieves import Pipeline
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.chunking.chonkie_ import Chonkie
import chonkie, tokenizers
# Build pipeline: Docling → Token-based chunking
pipeline = Pipeline([
Docling(export_format="markdown"),
Chonkie(chunker=chonkie.TokenChunker(tokenizers.Tokenizer.from_pretrained("gpt2")))
])
# Run on multiple documents
docs = [Doc(uri="file:///data/file1.pdf"), Doc(uri="file:///data/file2.pdf")]
processed = pipeline(docs)
for d in processed:
print(f"{d.uri} → {len(d.chunks)} chunks")
The Pipeline automatically respects each task's batch_size and caching semantics defined in sieves/tasks/core.py.
Summary
- Docling and Marker tasks in Sieves wrap third-party parsers to convert binary documents into structured text within the
Doc.textattribute. - Both tasks inherit from
sieves.tasks.core.Taskand follow the standard execution model: validate, delegate, store, yield. - Docling (
sieves/tasks/preprocessing/ingestion/docling_.py) supports multiple document formats (PDF, DOCX) with markdown/HTML export and automatic batch size tuning. - Marker (
sieves/tasks/preprocessing/ingestion/marker_.py) specializes in PDF conversion with optional image extraction intodoc.imagesand supports markdown/HTML/JSON output formats. - Install via
pip install sieves[ingestion]or directly installdocling/marker. - Both tasks integrate seamlessly into
sieves.Pipelineobjects for end-to-end document processing workflows.
Frequently Asked Questions
What file formats are supported by Docling and Marker in Sieves?
Docling supports PDF, DOCX, and other office formats through its DocumentConverter class, converting them to markdown or HTML. Marker focuses specifically on PDF files using PdfConverter or TableConverter, with output options including markdown, HTML, and JSON. Both tasks require documents to be provided as Doc objects with valid file URIs in the .uri attribute.
How do I extract images from PDFs using the Marker task?
Set the extract_images=True parameter when initializing the Marker task. When processing documents, Marker will extract embedded images and store them in the doc.images list of each processed Doc object. This occurs in the _call method of sieves/tasks/preprocessing/ingestion/marker_.py. Note that image extraction is specific to Marker; Docling handles images differently through its own document object model.
Can I filter which documents get processed by the ingestion tasks?
Yes, both Docling and Marker accept a condition callable parameter inherited from the base Task class in sieves/tasks/core.py. This function receives a Doc object and returns a boolean indicating whether to process it. For example, you can filter by file size, extension, or URI pattern. The task skips any documents where the condition returns False, allowing you to process only relevant files in large document collections.
What is the difference between the batch_size parameter in Docling versus Marker?
In Docling, the batch_size parameter automatically adjusts the underlying library's performance settings by modifying docling.datamodel.settings.settings.perf.doc_batch_size to optimize throughput for the specified batch size. In Marker, the batch_size parameter is handled by the base Task class and simply controls how many documents are passed to the task's _call method at once, without modifying internal Marker settings. Both respect the standard Sieves task execution model defined in sieves/tasks/core.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →