How to Track Token Usage and Input/Output Counts in Sieves

Sieves automatically records token consumption for every model call at per-chunk, per-task, and per-document granularities, storing the data in each Doc object's meta dictionary when include_meta=True (the default).

Tracking token usage is essential for cost management and performance optimization in AI pipelines. In the mantisai/sieves library, every predictive task captures input and output token counts without requiring manual instrumentation. This article explains how to access these metrics using the built-in observability features exposed through the Doc.meta attribute.

How Token Usage Tracking Works in Sieves

The library implements a three-tier tracking system that aggregates token data automatically during pipeline execution. When a predictive task processes a document, the framework captures usage at the chunk level, rolls it up to the task level, and maintains a running total at the document level.

The TokenUsage Model and Model Wrappers

At the lowest level, each model wrapper returns a TokenUsage object containing input_tokens and output_tokens. This model is defined in sieves/model_wrappers/types.py and is instantiated by wrapper implementations such as OutlinesModelWrapper (in outlines_.py) and HuggingFace wrappers (in huggingface_.py).

When a backend supplies native counts (like DSPy or LangChain), the wrapper extracts them directly from response metadata. Otherwise, the base ModelWrapper class in sieves/model_wrappers/core.py provides a _count_tokens method that uses a tokenizer or tiktoken to estimate counts:

from sieves.model_wrappers.core import ModelWrapper, TokenUsage

class MyWrapper(ModelWrapper):
    def _infer(self, ...):
        # ... run model and get raw_output ...

        
        # Estimate token counts with the shared helper

        tokenizer = self._get_tokenizer()
        input_tokens = self._count_tokens(prompt, tokenizer)
        output_tokens = self._count_tokens(str(raw_output), tokenizer)
        
        usage = TokenUsage(input_tokens=input_tokens, output_tokens=output_tokens)
        return [(parsed_result, raw_output, usage)]

Aggregation in PredictiveTask

The PredictiveTask class in sieves/tasks/predictive/core.py contains the _integrate_usage method, which handles the aggregation logic. This method receives the list of TokenUsage objects per chunk, maps chunk offsets to their parent documents, and aggregates the counts:


# Per-task usage stored under doc.meta[task_id]['usage']

doc.meta[self._task_id] = {
    "raw": raw_outputs[start:end],
    "usage": {
        "input_tokens": task_input_tokens,
        "output_tokens": task_output_tokens,
        "chunks": [u.model_dump() for u in doc_usages],
    },
}

# Top-level cumulative usage stored under doc.meta['usage']

usage = doc.meta.get("usage", {"input_tokens": None, "output_tokens": None})
usage["input_tokens"] = (usage["input_tokens"] or 0) + task_input_tokens
usage["output_tokens"] = (usage["output_tokens"] or 0) + task_output_tokens
doc.meta["usage"] = usage

Accessing Token Counts at Different Granularities

Sieves stores token data in the meta dictionary of each Doc object, organized by scope.

Per-Document Totals

To retrieve cumulative token consumption across all tasks in a pipeline, access the top-level usage key:

total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")

These values represent the sum of all input and output tokens consumed by every predictive task that processed the document.

Per-Task Breakdown

For task-specific metrics, access the task ID key within meta. Each predictive task stores its usage under its identifier (e.g., "Classification" or "SentimentAnalysis"):

cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")

Per-Chunk Details

To inspect token consumption for individual chunks (useful for debugging long documents), iterate over the chunks list:

for i, chunk in enumerate(cls_usage["chunks"]):
    print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")

Practical Implementation Examples

Inspecting Usage After Classification

Here is a complete example showing how to track token usage when running a classification task with the Outlines wrapper:

from sieves.tasks import Classification
from sieves.model_wrappers import OutlinesModelWrapper

# Create a classification task (include_meta is True by default)

task = Classification(
    labels=["science", "politics"], 
    model=OutlinesModelWrapper(model_name="gpt-4o-mini")
)

# Run the task on a list of Docs

docs = list(task(docs))

# Examine the first document's token statistics

doc = docs[0]

# Total tokens for the whole document (all tasks)

total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")

# Task-specific usage

cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")

# Per-chunk breakdown

for i, chunk in enumerate(cls_usage["chunks"]):
    print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")

Source: usage walkthrough in [docs/guides/observability.md](https://github.com/mantisai/sieves/blob/main/docs/guides/observability.md).

Summarizing Usage Across a Pipeline

When running multiple tasks in a pipeline, aggregate totals across all documents to generate cost reports:

from sieves.pipeline import Pipeline
from sieves.tasks import Classification, SentimentAnalysis

# Build a pipeline with multiple predictive tasks

pipeline = Pipeline(
    [Classification(labels=["a","b"]), SentimentAnalysis()], 
    use_cache=False
)

# Execute on a collection of Docs

docs = list(pipeline(docs))

# Aggregate totals across all documents

total_input = sum(doc.meta["usage"]["input_tokens"] or 0 for doc in docs)
total_output = sum(doc.meta["usage"]["output_tokens"] or 0 for doc in docs)

print(f"Pipeline consumed {total_input} input tokens and {total_output} output tokens.")

Summary

  • Automatic collection: Sieves captures token counts via TokenUsage objects returned by model wrappers in sieves/model_wrappers/core.py, using either native backend metadata or the _count_tokens tokenizer helper.
  • Three-tier storage: Per-chunk data resides in doc.meta[task_id]['usage']['chunks'], per-task totals in doc.meta[task_id]['usage'], and per-document totals in doc.meta['usage'].
  • Zero configuration: Tracking is enabled by default through include_meta=True in predictive tasks, with aggregation handled by PredictiveTask._integrate_usage in sieves/tasks/predictive/core.py.
  • Universal support: All built-in wrappers (Outlines, HuggingFace, GLiNER) implement the TokenUsage interface, and custom wrappers can leverage the base class counting utilities.

Frequently Asked Questions

Does token tracking work with all model wrappers in Sieves?

Yes. All built-in wrappers—including those for Outlines, Hugging Face Transformers, GLiNER, DSPy, and LangChain—return TokenUsage objects defined in sieves/model_wrappers/types.py. If a backend provides native token counts (like DSPy or LangChain), the wrapper extracts them directly; otherwise, it falls back to the _count_tokens method in the base ModelWrapper class.

How do I disable token usage tracking to reduce overhead?

Token tracking occurs only when include_meta=True (the default) on a predictive task. To disable it, instantiate the task with include_meta=False:

task = Classification(labels=["a", "b"], include_meta=False)

This prevents the _integrate_usage method from populating doc.meta with token data, slightly reducing memory overhead and processing time for high-volume pipelines.

Can I implement custom token counting for a new model wrapper?

Yes. When creating a custom ModelWrapper, implement the _infer method to return a list of tuples containing (parsed_result, raw_output, TokenUsage). Use the inherited _count_tokens method from sieves/model_wrappers/core.py to tokenize prompts and outputs if the underlying API does not provide native counts. The PredictiveTask infrastructure will automatically aggregate your custom counts into the document metadata.

What happens if a model backend doesn't provide native token counts?

The framework gracefully handles missing native counts by using the generic _count_tokens helper in the base wrapper class. This method attempts to load an appropriate tokenizer (such as tiktoken for OpenAI models or the model's own tokenizer for Hugging Face models) to estimate token counts. If tokenization fails or no tokenizer is available, the TokenUsage fields remain None, ensuring the pipeline continues executing without raising errors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →