# How to Use Observability and Logging with Loguru in Sieves Pipelines

> Master Loguru logging and observability in Sieves pipelines. Capture parsing errors and LLM outputs for better debugging and analysis. Learn how to integrate seamlessly.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Configure Loguru with custom handlers before importing Sieves to capture Docling parsing errors, then access token usage and raw LLM outputs via the `doc.meta` dictionary after pipeline execution.**

Sieves is an open-source Python library for building modular document processing pipelines. To monitor LLM token consumption, track raw model responses, and capture ingestion failures, the library stores observability data directly in document metadata while using Loguru for error logging in specific preprocessing tasks.

## Where Sieves Emits Logs

Only the **Docling ingestion task** uses Loguru directly. In [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py), the parser logs file processing failures at the ERROR level:

```python
from loguru import logger

# Inside the Docling task's processing logic

logger.error(f"Failed to parse file {doc.uri}: {str(e)}")

```

All other observability data—including token counts and model outputs—is stored silently in the `Doc.meta` dictionary rather than emitted as log messages. This design keeps log output clean while preserving rich execution metrics for programmatic access.

## Token Usage and Raw Output Tracking

The `PredictiveTask` class in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) (lines 16-38) automatically aggregates per-chunk LLM usage statistics after each batch. The `_integrate_usage` method injects this data into document metadata:

- **`doc.meta['usage']`** – Cumulative input and output token counts across **all** tasks in the pipeline
- **`doc.meta[task_id]['usage']`** – Task-specific token breakdowns including per-chunk details
- **`doc.meta[task_id]['raw']`** – List of raw LLM responses for each chunk (available when `include_meta=True`)

The aggregation logic sums token counts across document chunks and updates both task-level and pipeline-level summaries:

```python

# Simplified from sieves/tasks/predictive/core.py

task_input_tokens = sum(u.input_tokens for u in doc_usages if u.input_tokens)
task_output_tokens = sum(u.output_tokens for u in doc_usages if u.output_tokens)

# Store task-specific data

doc.meta[self._task_id] = {
    "raw": raw_outputs,
    "usage": {
        "input_tokens": task_input_tokens,
        "output_tokens": task_output_tokens,
        "chunks": [u.model_dump() for u in doc_usages],
    },
}

# Update cumulative totals

usage = doc.meta.get("usage", {"input_tokens": 0, "output_tokens": 0})
usage["input_tokens"] += task_input_tokens
usage["output_tokens"] += task_output_tokens
doc.meta["usage"] = usage

```

## Configuring Loguru for Sieves Pipelines

Because the Docling module imports Loguru at import time, you must configure Loguru **before** importing any Sieves components. This ensures all library calls respect your handlers, formatting, and level settings.

```python

# observability_setup.py

from loguru import logger
import sys
from pathlib import Path

# Remove default handler and configure custom sinks

logger.remove()
logger.add(
    sys.stderr,
    level="INFO",
    format="<green>{time:YYYY-MM-DD HH:mm:ss}</green> | <level>{level}</level> | {message}",
)
logger.add(
    Path("pipeline.log"),
    rotation="10 MB",
    retention="7 days",
    level="ERROR",
    enqueue=True,  # Thread-safe for concurrent tasks

)

# Now safe to import Sieves

from sieves import Pipeline, Doc
from sieves.tasks.preprocessing.ingestion.docling_ import Docling
from sieves.tasks.predictive.classification.core import Classification
from sieves.model_wrappers import ModelType, ModelSettings, Model

# Build pipeline with metadata tracking enabled

model = Model(type=ModelType.outlines, name="gpt-4o-mini")
settings = ModelSettings(inference_mode="text")

pipeline = Pipeline(
    Docling(),
    Classification(
        model=model,
        model_settings=settings,
        task_id="classifier",
        include_meta=True  # Required to capture raw outputs

    )
)

```

**Critical configuration details:**
- Call `logger.remove()` first to eliminate the default stdout handler
- Set `enqueue=True` on file handlers for thread-safe logging across parallel task execution
- Use separate handlers for console (INFO level) and file (ERROR level) to isolate Docling parsing failures

## Accessing Observability Data After Execution

Once the pipeline completes, query the `Doc` objects to retrieve execution metrics. The metadata structure provides both aggregate and granular visibility:

```python

# Run pipeline

docs = [Doc(uri=Path("document.pdf"))]
pipeline(docs)

# Extract observability data

for doc in docs:
    # Cumulative token usage across all tasks

    total_usage = doc.meta.get("usage", {})
    print(f"Total input tokens: {total_usage.get('input_tokens')}")
    print(f"Total output tokens: {total_usage.get('output_tokens')}")
    
    # Task-specific metrics (using the task_id defined in pipeline)

    task_data = doc.meta.get("classifier", {})
    if task_data:
        print(f"Task input tokens: {task_data['usage']['input_tokens']}")
        print(f"Number of chunks processed: {len(task_data['usage']['chunks'])}")
        
        # Raw LLM responses for debugging or audit trails

        raw_outputs = task_data.get("raw", [])
        for i, response in enumerate(raw_outputs):
            print(f"Chunk {i} response: {response[:100]}...")

```

## Extending Logging in Custom Tasks

For pipelines requiring additional instrumentation, inject Loguru calls into custom `Task` subclasses. Because Loguru is already configured globally, these messages automatically flow to your configured handlers:

```python
from loguru import logger
from sieves.tasks.core import Task
from typing import Iterable
from sieves.data.doc import Doc

class InstrumentedTask(Task):
    def _call(self, docs: Iterable[Doc]) -> Iterable[Doc]:
        doc_list = list(docs)
        logger.info(f"Starting {self.__class__.__name__} on {len(doc_list)} documents")
        
        # Processing logic here

        processed = 0
        for doc in doc_list:
            processed += 1
        
        logger.success(f"Completed {self.__class__.__name__}, processed {processed} docs")
        return doc_list

```

This pattern enables granular performance tracking and debugging without modifying the core Sieves library.

## Summary

- **Configure Loguru before importing Sieves** to ensure the Docling error logger respects your handlers and formatting
- **Docling parsing errors** are captured at ERROR level in [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py) and routed to your configured sinks
- **Token usage aggregates** in `doc.meta['usage']` with per-task breakdowns available under `doc.meta[task_id]['usage']`
- **Raw LLM responses** are stored when `include_meta=True`, accessible via `doc.meta[task_id]['raw']` for debugging or audit trails
- **Extend observability** by adding Loguru calls to custom `Task` subclasses; all logs respect the global configuration

## Frequently Asked Questions

### When should I configure Loguru in my Sieves script?

Configure Loguru **before** any Sieves import statement executes. The [`sieves/tasks/preprocessing/ingestion/docling_.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/preprocessing/ingestion/docling_.py) module imports `from loguru import logger` at the top level, so configuring handlers after this import means the default logger will already be initialized. Place your `logger.remove()` and `logger.add()` calls at the very start of your script, then import Sieves components.

### Where are token usage statistics stored after running a pipeline?

Cumulative token counts across all tasks reside in `doc.meta['usage']` with keys `input_tokens` and `output_tokens`. For task-specific breakdowns including per-chunk details, access `doc.meta[task_id]['usage']` where `task_id` matches the string identifier you passed to the task constructor.

### How do I access raw LLM responses for debugging?

Raw responses are available only when you instantiate predictive tasks with `include_meta=True`. After execution, retrieve them via `doc.meta[task_id]['raw']`, which returns a list of response strings corresponding to each document chunk processed by that task.

### Can I add custom logging to my own Sieves tasks?

Yes. Import `logger` from Loguru inside your custom `Task` subclass and emit logs at any level (INFO, DEBUG, ERROR). Because Loguru uses a global configuration, these messages automatically respect the handlers, formatting, and rotation policies you established at script startup.