How to Track Token Usage and Input/Output Counts in Sieves
Sieves automatically records token consumption for every model call at per-chunk, per-task, and per-document granularities, storing the data in each Doc object's meta dictionary when include_meta=True (the default).
Tracking token usage is essential for cost management and performance optimization in AI pipelines. In the mantisai/sieves library, every predictive task captures input and output token counts without requiring manual instrumentation. This article explains how to access these metrics using the built-in observability features exposed through the Doc.meta attribute.
How Token Usage Tracking Works in Sieves
The library implements a three-tier tracking system that aggregates token data automatically during pipeline execution. When a predictive task processes a document, the framework captures usage at the chunk level, rolls it up to the task level, and maintains a running total at the document level.
The TokenUsage Model and Model Wrappers
At the lowest level, each model wrapper returns a TokenUsage object containing input_tokens and output_tokens. This model is defined in sieves/model_wrappers/types.py and is instantiated by wrapper implementations such as OutlinesModelWrapper (in outlines_.py) and HuggingFace wrappers (in huggingface_.py).
When a backend supplies native counts (like DSPy or LangChain), the wrapper extracts them directly from response metadata. Otherwise, the base ModelWrapper class in sieves/model_wrappers/core.py provides a _count_tokens method that uses a tokenizer or tiktoken to estimate counts:
from sieves.model_wrappers.core import ModelWrapper, TokenUsage
class MyWrapper(ModelWrapper):
def _infer(self, ...):
# ... run model and get raw_output ...
# Estimate token counts with the shared helper
tokenizer = self._get_tokenizer()
input_tokens = self._count_tokens(prompt, tokenizer)
output_tokens = self._count_tokens(str(raw_output), tokenizer)
usage = TokenUsage(input_tokens=input_tokens, output_tokens=output_tokens)
return [(parsed_result, raw_output, usage)]
Aggregation in PredictiveTask
The PredictiveTask class in sieves/tasks/predictive/core.py contains the _integrate_usage method, which handles the aggregation logic. This method receives the list of TokenUsage objects per chunk, maps chunk offsets to their parent documents, and aggregates the counts:
# Per-task usage stored under doc.meta[task_id]['usage']
doc.meta[self._task_id] = {
"raw": raw_outputs[start:end],
"usage": {
"input_tokens": task_input_tokens,
"output_tokens": task_output_tokens,
"chunks": [u.model_dump() for u in doc_usages],
},
}
# Top-level cumulative usage stored under doc.meta['usage']
usage = doc.meta.get("usage", {"input_tokens": None, "output_tokens": None})
usage["input_tokens"] = (usage["input_tokens"] or 0) + task_input_tokens
usage["output_tokens"] = (usage["output_tokens"] or 0) + task_output_tokens
doc.meta["usage"] = usage
Accessing Token Counts at Different Granularities
Sieves stores token data in the meta dictionary of each Doc object, organized by scope.
Per-Document Totals
To retrieve cumulative token consumption across all tasks in a pipeline, access the top-level usage key:
total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")
These values represent the sum of all input and output tokens consumed by every predictive task that processed the document.
Per-Task Breakdown
For task-specific metrics, access the task ID key within meta. Each predictive task stores its usage under its identifier (e.g., "Classification" or "SentimentAnalysis"):
cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")
Per-Chunk Details
To inspect token consumption for individual chunks (useful for debugging long documents), iterate over the chunks list:
for i, chunk in enumerate(cls_usage["chunks"]):
print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")
Practical Implementation Examples
Inspecting Usage After Classification
Here is a complete example showing how to track token usage when running a classification task with the Outlines wrapper:
from sieves.tasks import Classification
from sieves.model_wrappers import OutlinesModelWrapper
# Create a classification task (include_meta is True by default)
task = Classification(
labels=["science", "politics"],
model=OutlinesModelWrapper(model_name="gpt-4o-mini")
)
# Run the task on a list of Docs
docs = list(task(docs))
# Examine the first document's token statistics
doc = docs[0]
# Total tokens for the whole document (all tasks)
total = doc.meta["usage"]
print(f"Total input tokens: {total['input_tokens']}, total output tokens: {total['output_tokens']}")
# Task-specific usage
cls_usage = doc.meta["Classification"]["usage"]
print(f"Classification task input tokens: {cls_usage['input_tokens']}")
print(f"Classification task output tokens: {cls_usage['output_tokens']}")
# Per-chunk breakdown
for i, chunk in enumerate(cls_usage["chunks"]):
print(f"Chunk {i}: {chunk['input_tokens']} in, {chunk['output_tokens']} out")
Source: usage walkthrough in [docs/guides/observability.md](https://github.com/mantisai/sieves/blob/main/docs/guides/observability.md).
Summarizing Usage Across a Pipeline
When running multiple tasks in a pipeline, aggregate totals across all documents to generate cost reports:
from sieves.pipeline import Pipeline
from sieves.tasks import Classification, SentimentAnalysis
# Build a pipeline with multiple predictive tasks
pipeline = Pipeline(
[Classification(labels=["a","b"]), SentimentAnalysis()],
use_cache=False
)
# Execute on a collection of Docs
docs = list(pipeline(docs))
# Aggregate totals across all documents
total_input = sum(doc.meta["usage"]["input_tokens"] or 0 for doc in docs)
total_output = sum(doc.meta["usage"]["output_tokens"] or 0 for doc in docs)
print(f"Pipeline consumed {total_input} input tokens and {total_output} output tokens.")
Summary
- Automatic collection: Sieves captures token counts via
TokenUsageobjects returned by model wrappers insieves/model_wrappers/core.py, using either native backend metadata or the_count_tokenstokenizer helper. - Three-tier storage: Per-chunk data resides in
doc.meta[task_id]['usage']['chunks'], per-task totals indoc.meta[task_id]['usage'], and per-document totals indoc.meta['usage']. - Zero configuration: Tracking is enabled by default through
include_meta=Truein predictive tasks, with aggregation handled byPredictiveTask._integrate_usageinsieves/tasks/predictive/core.py. - Universal support: All built-in wrappers (Outlines, HuggingFace, GLiNER) implement the
TokenUsageinterface, and custom wrappers can leverage the base class counting utilities.
Frequently Asked Questions
Does token tracking work with all model wrappers in Sieves?
Yes. All built-in wrappers—including those for Outlines, Hugging Face Transformers, GLiNER, DSPy, and LangChain—return TokenUsage objects defined in sieves/model_wrappers/types.py. If a backend provides native token counts (like DSPy or LangChain), the wrapper extracts them directly; otherwise, it falls back to the _count_tokens method in the base ModelWrapper class.
How do I disable token usage tracking to reduce overhead?
Token tracking occurs only when include_meta=True (the default) on a predictive task. To disable it, instantiate the task with include_meta=False:
task = Classification(labels=["a", "b"], include_meta=False)
This prevents the _integrate_usage method from populating doc.meta with token data, slightly reducing memory overhead and processing time for high-volume pipelines.
Can I implement custom token counting for a new model wrapper?
Yes. When creating a custom ModelWrapper, implement the _infer method to return a list of tuples containing (parsed_result, raw_output, TokenUsage). Use the inherited _count_tokens method from sieves/model_wrappers/core.py to tokenize prompts and outputs if the underlying API does not provide native counts. The PredictiveTask infrastructure will automatically aggregate your custom counts into the document metadata.
What happens if a model backend doesn't provide native token counts?
The framework gracefully handles missing native counts by using the generic _count_tokens helper in the base wrapper class. This method attempts to load an appropriate tokenizer (such as tiktoken for OpenAI models or the model's own tokenizer for Hugging Face models) to estimate token counts. If tokenization fails or no tokenizer is available, the TokenUsage fields remain None, ensuring the pipeline continues executing without raising errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →