# How to Track Token Usage and Input/Output Counts in Sieves

> Easily track token usage input and output counts in your Sieves projects. Learn how Sieves automatically records valuable data for every model call within your Doc objects.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Sieves automatically records token consumption for every model call at per-chunk, per-task, and per-document granularities, storing the data in each `Doc` object's `meta` dictionary when `include_meta=True` (the default).**

Tracking token usage is essential for cost management and performance optimization in AI pipelines. In the `mantisai/sieves` library, every predictive task captures input and output token counts without requiring manual instrumentation. This article explains how to access these metrics using the built-in observability features exposed through the `Doc.meta` attribute.

## How Token Usage Tracking Works in Sieves

The library implements a three-tier tracking system that aggregates token data automatically during pipeline execution. When a **predictive task** processes a document, the framework captures usage at the chunk level, rolls it up to the task level, and maintains a running total at the document level.

### The TokenUsage Model and Model Wrappers

At the lowest level, each **model wrapper** returns a `TokenUsage` object containing `input_tokens` and `output_tokens`. This model is defined in [`sieves/model_wrappers/types.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py) and is instantiated by wrapper implementations such as `OutlinesModelWrapper` (in [`outlines_.py`](https://github.com/mantisai/sieves/blob/main/outlines_.py)) and `HuggingFace` wrappers (in [`huggingface_.py`](https://github.com/mantisai/sieves/blob/main/huggingface_.py)).

When a backend supplies native counts (like DSPy or LangChain), the wrapper extracts them directly from response metadata. Otherwise, the base `ModelWrapper` class in [`sieves/model_wrappers/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/core.py) provides a `_count_tokens` method that uses a tokenizer or `tiktoken` to estimate counts:

```python
from sieves.model_wrappers.core import ModelWrapper, TokenUsage

class MyWrapper(ModelWrapper):
    def _infer(self, ...):
        # ... run model and get raw_output ...

        
        # Estimate token counts with the shared helper

        tokenizer = self._get_tokenizer()
        input_tokens = self._count_tokens(prompt, tokenizer)
        output_tokens = self._count_tokens(str(raw_output), tokenizer)
        
        usage = TokenUsage(input_tokens=input_tokens, output_tokens=output_tokens)
        return [(parsed_result, raw_output, usage)]

```

### Aggregation in PredictiveTask

The `PredictiveTask` class in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py) contains the `_integrate_usage` method, which handles the aggregation logic. This method receives the list of `TokenUsage` objects per chunk, maps chunk offsets to their parent documents, and aggregates the counts:

```python

# Per-task usage stored under doc.meta[task_id]['usage']

doc.meta[self._task_id] = {
    "raw": raw_outputs[start:end],
    "usage": {
        "input_tokens": task_input_tokens,
        "output_tokens": task_output_tokens,
        "chunks": [u.model_dump() for u in doc_usages],
    },
}

# Top-level cumulative usage stored under doc.meta['usage']

usage = doc.meta.get("usage", {"input_tokens": None, "output_tokens": None})
usage["input_tokens"] = (usage["input_tokens"] or 0) + task_input_tokens
usage["output_tokens"] = (usage["output_tokens"] or 0) + task_output_tokens
doc.meta["usage"] = usage

```

## Accessing Token Counts at Different Granularities

Sieves stores token data in the `meta` dictionary of each `Doc` object, organized by scope.

### Per-Document Totals

To retrieve cumulative token consumption across all tasks in a pipeline, access the top-level `usage` key:

```python
total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")

```

These values represent the sum of all input and output tokens consumed by every predictive task that processed the document.

### Per-Task Breakdown

For task-specific metrics, access the task ID key within `meta`. Each predictive task stores its usage under its identifier (e.g., `"Classification"` or `"SentimentAnalysis"`):

```python
cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")

```

### Per-Chunk Details

To inspect token consumption for individual chunks (useful for debugging long documents), iterate over the `chunks` list:

```python
for i, chunk in enumerate(cls_usage["chunks"]):
    print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")

```

## Practical Implementation Examples

### Inspecting Usage After Classification

Here is a complete example showing how to track token usage when running a classification task with the Outlines wrapper:

```python
from sieves.tasks import Classification
from sieves.model_wrappers import OutlinesModelWrapper

# Create a classification task (include_meta is True by default)

task = Classification(
    labels=["science", "politics"], 
    model=OutlinesModelWrapper(model_name="gpt-4o-mini")
)

# Run the task on a list of Docs

docs = list(task(docs))

# Examine the first document's token statistics

doc = docs[0]

# Total tokens for the whole document (all tasks)

total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")

# Task-specific usage

cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")

# Per-chunk breakdown

for i, chunk in enumerate(cls_usage["chunks"]):
    print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")

```

*Source: usage walkthrough in* [[`docs/guides/observability.md`](https://github.com/mantisai/sieves/blob/main/docs/guides/observability.md)](https://github.com/mantisai/sieves/blob/main/docs/guides/observability.md).

### Summarizing Usage Across a Pipeline

When running multiple tasks in a pipeline, aggregate totals across all documents to generate cost reports:

```python
from sieves.pipeline import Pipeline
from sieves.tasks import Classification, SentimentAnalysis

# Build a pipeline with multiple predictive tasks

pipeline = Pipeline(
    [Classification(labels=["a","b"]), SentimentAnalysis()], 
    use_cache=False
)

# Execute on a collection of Docs

docs = list(pipeline(docs))

# Aggregate totals across all documents

total_input = sum(doc.meta["usage"]["input_tokens"] or 0 for doc in docs)
total_output = sum(doc.meta["usage"]["output_tokens"] or 0 for doc in docs)

print(f"Pipeline consumed {total_input} input tokens and {total_output} output tokens.")

```

## Summary

- **Automatic collection**: Sieves captures token counts via `TokenUsage` objects returned by model wrappers in [`sieves/model_wrappers/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/core.py), using either native backend metadata or the `_count_tokens` tokenizer helper.
- **Three-tier storage**: Per-chunk data resides in `doc.meta[task_id]['usage']['chunks']`, per-task totals in `doc.meta[task_id]['usage']`, and per-document totals in `doc.meta['usage']`.
- **Zero configuration**: Tracking is enabled by default through `include_meta=True` in predictive tasks, with aggregation handled by `PredictiveTask._integrate_usage` in [`sieves/tasks/predictive/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/core.py).
- **Universal support**: All built-in wrappers (Outlines, HuggingFace, GLiNER) implement the `TokenUsage` interface, and custom wrappers can leverage the base class counting utilities.

## Frequently Asked Questions

### Does token tracking work with all model wrappers in Sieves?

Yes. All built-in wrappers—including those for Outlines, Hugging Face Transformers, GLiNER, DSPy, and LangChain—return `TokenUsage` objects defined in [`sieves/model_wrappers/types.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py). If a backend provides native token counts (like DSPy or LangChain), the wrapper extracts them directly; otherwise, it falls back to the `_count_tokens` method in the base `ModelWrapper` class.

### How do I disable token usage tracking to reduce overhead?

Token tracking occurs only when `include_meta=True` (the default) on a predictive task. To disable it, instantiate the task with `include_meta=False`:

```python
task = Classification(labels=["a", "b"], include_meta=False)

```

This prevents the `_integrate_usage` method from populating `doc.meta` with token data, slightly reducing memory overhead and processing time for high-volume pipelines.

### Can I implement custom token counting for a new model wrapper?

Yes. When creating a custom `ModelWrapper`, implement the `_infer` method to return a list of tuples containing `(parsed_result, raw_output, TokenUsage)`. Use the inherited `_count_tokens` method from [`sieves/model_wrappers/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/core.py) to tokenize prompts and outputs if the underlying API does not provide native counts. The `PredictiveTask` infrastructure will automatically aggregate your custom counts into the document metadata.

### What happens if a model backend doesn't provide native token counts?

The framework gracefully handles missing native counts by using the generic `_count_tokens` helper in the base wrapper class. This method attempts to load an appropriate tokenizer (such as `tiktoken` for OpenAI models or the model's own tokenizer for Hugging Face models) to estimate token counts. If tokenization fails or no tokenizer is available, the `TokenUsage` fields remain `None`, ensuring the pipeline continues executing without raising errors.