# How to Perform Document Ingestion with Docling and Marker in Sieves

> Learn to ingest documents with Docling and Marker in the Sieves repository. Easily convert PDFs to plain text and extract content for streamlined processing.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Use the `Docling` or `Marker` tasks in Sieves to convert PDFs and documents into plain text by passing `Doc` objects with file URIs, which automatically extract content into `doc.text` for downstream processing.**

Sieves provides dedicated pre-processing ingestion tasks that wrap **Docling** and **Marker**, two popular open-source document parsing libraries. These tasks handle the heavy lifting of converting binary documents into structured text within the `sieves` framework, enabling seamless integration with chunking, embedding, and predictive modeling pipelines.

## Understanding Document Ingestion in Sieves

Both `Docling` and `Marker` inherit from `sieves.tasks.core.Task` and follow a standardized five-step execution model defined in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py):

1. **Receive** an iterable of `sieves.data.Doc` objects, each requiring a `.uri` attribute pointing to the source file.
2. **Validate** incoming documents to ensure URIs are present and warn if existing `.text` will be overwritten.
3. **Delegate** parsing to the external library (`docling.document_converter.DocumentConverter` or `marker.converters.pdf.PdfConverter`).
4. **Store** extracted content in `doc.text` and optionally attach metadata or images.
5. **Yield** enriched `Doc` objects for downstream consumption.

### Docling vs. Marker: Key Differences

| Feature | Docling Ingestion | Marker Ingestion |
|---------|------------------|------------------|
| **Core Class** | `Docling(Task)` in [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py) | `Marker(Task)` in [`sieves/tasks/preprocessing/ingestion/marker_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/marker_.py) |
| **Supported Formats** | PDF, DOCX, and other formats via `DocumentConverter` | PDF-focused via `PdfConverter`/`TableConverter` |
| **Export Formats** | `markdown` (default), `html` | `markdown` (default), `html`, `json` |
| **Batch Handling** | Auto-adjusts `docling.datamodel.settings.settings.perf.doc_batch_size` based on task `batch_size` | No internal batch tuning; relies on task iteration |
| **Image Extraction** | Available via Docling's native capabilities | Native support via `extract_images=True` parameter |
| **Metadata Storage** | Stores raw `ParsedResource` in `doc.meta[task_id]` when `include_meta=True` | Stores `Converter` instance in serialized state; images attach to `doc.images` |

## Installation and Setup

Both tasks are available as optional extras. Install them using the `ingestion` extra or directly install the underlying libraries:

```bash

# Install via sieves extras

pip install sieves[ingestion]

# Or install parsers directly

pip install docling
pip install marker

```

Import the tasks from their respective modules:

```python
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.ingestion.marker_ import Marker
from sieves.data import Doc

```

## Performing Document Ingestion with Docling

Docling converts documents into markdown or HTML format using IBM's `DocumentConverter`. The task automatically handles batch processing and integrates with Docling's performance settings.

### Basic Docling Ingestion

Create a `Doc` with a file URI and process it through the `Docling` task:

```python
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling

# Create a Doc pointing at a local file

doc = Doc(uri="file:///path/to/report.pdf")

# Initialize the Docling task with markdown output (default)

docling_task = Docling(export_format="markdown")

# Run the task

processed_docs = list(docling_task([doc]))

# Access extracted text

print(processed_docs[0].text)  # Markdown string

# Access metadata if include_meta=True

# print(processed_docs[0].meta)  # {'docling': ParsedResource(...)}

```

The constructor and `_call` implementation are defined in lines 23-84 of [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py).

### Advanced Docling with Batching and Filtering

Use the `batch_size` parameter to control throughput and `condition` to filter documents before processing:

```python
from pathlib import Path
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling

# Build a collection of Docs from a directory

docs = [Doc(uri=str(p)) for p in Path("data/reports").glob("*.pdf")]

# Filter: only process PDFs larger than 1 MiB

def large_file(doc: Doc) -> bool:
    return Path(doc.uri).stat().st_size > 1_048_576

# Configure task with HTML output and custom batching

docling = Docling(
    export_format="html",
    batch_size=10,           # Process in batches of 10

    condition=large_file,      # Skip tiny files

    include_meta=True,
)

processed = list(docling(docs))
print(f"Parsed {len(processed)} large documents.")

```

The batch size is applied in `__init__` (lines 44-48) and the `condition` argument is passed to the base `Task` (line 40) in [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py).

## Performing Document Ingestion with Marker

Marker specializes in PDF-to-markdown conversion with optional image extraction. It supports multiple output formats including markdown, HTML, and JSON.

### Marker with Image Extraction

Process PDFs and extract embedded images into the `Doc.images` attribute:

```python
from sieves.data import Doc
from sieves.tasks.preprocessing.ingestion.marker_ import Marker

doc = Doc(uri="file:///path/to/paper.pdf")

marker = Marker(
    export_format="markdown",
    extract_images=True,   # Store extracted images in doc.images

    include_meta=False,
)

# Marker yields an iterator; take the first result

doc = next(marker([doc]))

print(doc.text)            # Markdown representation of the PDF

print(f"Found {len(doc.images)} images.")  # Images extracted by Marker

```

Image extraction is performed in the `_call` method at lines 9-12 of the loop in [`sieves/tasks/preprocessing/ingestion/marker_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/marker_.py).

## Integrating Ingestion into a Pipeline

Ingestion tasks work seamlessly within `sieves.Pipeline` objects, allowing you to chain document parsing with chunking and predictive modeling:

```python
from sieves import Pipeline
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.preprocessing.chunking.chonkie_ import Chonkie
import chonkie, tokenizers

# Build pipeline: Docling → Token-based chunking

pipeline = Pipeline([
    Docling(export_format="markdown"),
    Chonkie(chunker=chonkie.TokenChunker(tokenizers.Tokenizer.from_pretrained("gpt2")))
])

# Run on multiple documents

docs = [Doc(uri="file:///data/file1.pdf"), Doc(uri="file:///data/file2.pdf")]
processed = pipeline(docs)

for d in processed:
    print(f"{d.uri} → {len(d.chunks)} chunks")

```

The `Pipeline` automatically respects each task's `batch_size` and caching semantics defined in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py).

## Summary

- **Docling and Marker tasks** in Sieves wrap third-party parsers to convert binary documents into structured text within the `Doc.text` attribute.
- Both tasks inherit from `sieves.tasks.core.Task` and follow the standard execution model: validate, delegate, store, yield.
- **Docling** ([`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py)) supports multiple document formats (PDF, DOCX) with markdown/HTML export and automatic batch size tuning.
- **Marker** ([`sieves/tasks/preprocessing/ingestion/marker_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/marker_.py)) specializes in PDF conversion with optional image extraction into `doc.images` and supports markdown/HTML/JSON output formats.
- Install via `pip install sieves[ingestion]` or directly install `docling`/`marker`.
- Both tasks integrate seamlessly into `sieves.Pipeline` objects for end-to-end document processing workflows.

## Frequently Asked Questions

### What file formats are supported by Docling and Marker in Sieves?

**Docling** supports PDF, DOCX, and other office formats through its `DocumentConverter` class, converting them to markdown or HTML. **Marker** focuses specifically on PDF files using `PdfConverter` or `TableConverter`, with output options including markdown, HTML, and JSON. Both tasks require documents to be provided as `Doc` objects with valid file URIs in the `.uri` attribute.

### How do I extract images from PDFs using the Marker task?

Set the `extract_images=True` parameter when initializing the `Marker` task. When processing documents, Marker will extract embedded images and store them in the `doc.images` list of each processed `Doc` object. This occurs in the `_call` method of [`sieves/tasks/preprocessing/ingestion/marker_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/marker_.py). Note that image extraction is specific to Marker; Docling handles images differently through its own document object model.

### Can I filter which documents get processed by the ingestion tasks?

Yes, both `Docling` and `Marker` accept a `condition` callable parameter inherited from the base `Task` class in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py). This function receives a `Doc` object and returns a boolean indicating whether to process it. For example, you can filter by file size, extension, or URI pattern. The task skips any documents where the condition returns `False`, allowing you to process only relevant files in large document collections.

### What is the difference between the batch_size parameter in Docling versus Marker?

In **Docling**, the `batch_size` parameter automatically adjusts the underlying library's performance settings by modifying `docling.datamodel.settings.settings.perf.doc_batch_size` to optimize throughput for the specified batch size. In **Marker**, the `batch_size` parameter is handled by the base `Task` class and simply controls how many documents are passed to the task's `_call` method at once, without modifying internal Marker settings. Both respect the standard Sieves task execution model defined in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py).