How to Access Raw Model Outputs and Debugging Meta Information in Sieves

Enable include_meta=True when initializing any predictive task in Sieves to capture raw LLM responses, token usage statistics, and per-chunk debugging data in the Doc.meta dictionary.

The Sieves library provides transparent access to model internals for debugging and cost monitoring. When you access raw model outputs and debugging meta information in Sieves, you gain visibility into exactly what the LLM returns before any post-processing occurs, alongside detailed token consumption metrics.

Enabling Meta Information Capture

Sieves stores predictive task outputs in the Doc.meta dictionary. Two mechanisms control what debugging information is preserved during pipeline execution.

Task-Level Configuration

Set include_meta=True during task instantiation to enable raw output capture. This flag is defined in the base Task class at sieves/tasks/core.py and inherited by all predictive tasks.

from sieves.tasks.predictive.classification import Classification

task = Classification(
    model="gpt-4o-mini",
    task_id="sentiment",
    include_meta=True,  # Enable raw output and usage tracking

    batch_size=-1,
)

How Meta Information is Stored

When include_meta=True, the _integrate_usage method in sieves/tasks/predictive/core.py (lines 16-27) populates doc.meta with a structured dictionary containing:

  • Raw outputs: Complete LLM responses per chunk
  • Usage statistics: Input/output token counts and per-chunk breakdowns
  • Task identification: Organized by task_id for multi-task pipelines

Accessing Raw LLM Responses and Token Usage

Once meta capture is enabled, you can extract debugging information at the document level after pipeline execution.

Retrieving Raw Model Outputs

Access the complete, unprocessed LLM responses through the task-specific key in Doc.meta:

from sieves.data import Doc
from sieves.pipeline import Pipeline
from sieves.tasks.predictive.translation import Translation

# Initialize task with meta capture

translate = Translation(
    model="gpt-4o-mini",
    task_id="translate",
    include_meta=True,
    batch_size=-1,
)

# Execute pipeline

pipe = Pipeline([translate])
doc = Doc(text="Hello world!", id="example-1")
result_docs = list(pipe([doc]))

# Access raw LLM responses

raw_responses = result_docs[0].meta["translate"]["raw"]
print("Raw model outputs:", raw_responses)

Source references: Task initialization in sieves/tasks/core.py (lines 24-33); meta integration in sieves/tasks/predictive/core.py (lines 16-27).

Examining Token Usage Statistics

Monitor API costs and consumption patterns through the usage statistics stored alongside raw outputs:


# Continuing from previous example

doc = result_docs[0]

# Per-task usage breakdown

task_usage = doc.meta["translate"]["usage"]
print("Input tokens:", task_usage["input_tokens"])
print("Output tokens:", task_usage["output_tokens"])

# Pipeline-wide cumulative usage

global_usage = doc.meta["usage"]
print("Total pipeline tokens:", global_usage)

The usage dictionary includes:

  • input_tokens: Total tokens sent to the model
  • output_tokens: Total tokens received from the model
  • chunks: Detailed per-chunk usage statistics (serialized via model_dump())

Multi-Chunk Document Handling

For long documents processed in multiple chunks (e.g., via Chonkie integration), Sieves preserves raw outputs for each chunk separately:

from sieves.tasks.predictive.ner import NER

ner = NER(
    model="gpt-4o-mini",
    task_id="ner",
    include_meta=True,
    batch_size=-1,
)

pipe = Pipeline([ner])

# Long text that will be chunked automatically

doc = Doc(
    text="Alice went to Paris. Bob stayed in Berlin. Charlie visited London.",
    id="multi-chunk-example",
)

result = list(pipe([doc]))[0]

# Each chunk has its own raw output entry

print("Number of chunks processed:", len(result.meta["ner"]["raw"]))
print("Raw outputs per chunk:", result.meta["ner"]["raw"])

Key Implementation Files

Understanding the source structure helps when debugging or extending meta capture functionality:

File Purpose
sieves/tasks/core.py Base Task class; defines include_meta flag and serialization logic
sieves/tasks/predictive/core.py Implements _integrate_usage method that stores raw outputs and token usage in Doc.meta
sieves/data/doc.py Doc model containing the meta dictionary that stores debugging information
sieves/pipeline.py Orchestrates task execution and manages document flow through the processing chain

Summary

  • Enable meta capture by setting include_meta=True when initializing any predictive task in sieves/tasks/predictive/.
  • Access raw outputs through doc.meta[task_id]["raw"] after pipeline execution to see unprocessed LLM responses.
  • Monitor token usage via doc.meta[task_id]["usage"] for per-task costs and doc.meta["usage"] for pipeline-wide totals.
  • Debug multi-chunk documents by examining the list of raw outputs stored for each processed chunk.

Frequently Asked Questions

How do I enable debugging mode in Sieves to see what the LLM returns?

Set include_meta=True when creating any predictive task such as Classification, NER, or Translation. This flag, defined in sieves/tasks/core.py, instructs the task to store raw LLM responses and token usage statistics in the Doc.meta dictionary after pipeline execution.

Where are token usage statistics stored in Sieves documents?

Token usage data is stored in two locations within Doc.meta: the task-specific entry at doc.meta[task_id]["usage"] contains input/output token counts and per-chunk breakdowns for that specific task, while doc.meta["usage"] maintains cumulative totals across all tasks in the pipeline. This dual structure allows both granular debugging and high-level cost monitoring.

Can I access raw outputs for documents processed in multiple chunks?

Yes. When Sieves processes long documents in multiple chunks (automatically handled by the chunking engine), the include_meta=True setting preserves raw LLM responses for each individual chunk. Access these through doc.meta[task_id]["raw"], which returns a list where each element corresponds to one chunk's raw model output. This is particularly useful for debugging inconsistent results across document sections.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →