How to Handle PII Masking and Anonymization in Sieves: A Complete Developer Guide

Sieves provides a dedicated PIIMasking task that automatically detects, masks, and scores personally identifiable information using pluggable model backends including DSPy, LangChain, and Outlines.

The mantisai/sieves library offers a production-ready solution for PII masking and anonymization through its predictive task framework. The PIIMasking class integrates seamlessly into the Pipeline architecture, providing consistent entity detection and text redaction across multiple LLM providers while maintaining strict separation between data models, orchestration logic, and model-specific implementations.

Architecture of the PII Masking System

The PII masking implementation follows a three-layer architecture that separates concerns between data validation, task orchestration, and model interaction.

Data Models and Schemas

At the foundation lies the schema layer in sieves/tasks/predictive/schemas/pii_masking.py, which defines three critical Pydantic models:

  • PIIEntity – A frozen model describing a single PII span with entity_type, text, and an optional score attribute for confidence levels.
  • Result – The unified task output containing the masked_text version and a list of PIIEntity objects representing detected entities.
  • FewshotExample – A wrapper for few-shot prompting that exposes text, masked_text, and pii_entities fields.

These schemas enforce type safety across all model-wrapper bridges, ensuring that DSPy, LangChain, and Outlines implementations return identically structured results regardless of underlying LLM differences.

Task Orchestration Layer

The core task logic resides in sieves/tasks/predictive/pii_masking/core.py, where the PIIMasking class subclasses PredictiveTask to provide:

  • Flexible configuration via constructor parameters including model, pii_types (accepting lists, dictionaries, or None for defaults), overwrite to replace original text, custom prompt_instructions, fewshot_examples, and ModelSettings for inference mode overrides.
  • Prompt signature defined by returning the Result model from prompt_signature, standardizing the expected output structure.
  • Corpus-level evaluation through _compute_metrics, which calculates micro-F1 scores based on exact matches of (entity_type, text) tuples, with a specialized _evaluate_dspy_example for DSPy-specific per-example evaluation.
  • Dataset export via to_hf_dataset, creating a Hugging Face datasets.Dataset with text and masked_text columns using the version string from Config.
  • Serialization support through _state, which preserves the pii_types configuration (including dictionary descriptions when provided) for pipeline persistence.

The core intentionally contains no model-specific logic, delegating all LLM interactions to the bridge layer via _init_bridge.

Bridge Abstraction for Model Backends

The bridge layer in sieves/tasks/predictive/pii_masking/bridges.py abstracts model-specific implementations behind the PIIMaskingBridge base class:

  • Placeholder management – The base bridge prepares the [MASKED] placeholder used to replace sensitive text spans.
  • PII type parsing – Handles pii_types as either a list or dictionary, with dictionary entries generating human-readable descriptions injected into prompts via an XML-style <pii_type_descriptions> block.
  • Runtime type generation – Creates a runtime PIIEntity class with a literal union of allowed entity types for strict validation.

Concrete implementations include:

  • DSPyPIIMasking – Constructs DSPy PromptSignature/Result pairs, extracts entities via res.pii_entities, and consolidates chunked results using MultiEntityConsolidation.
  • PydanticPIIMasking – Powers both LangChain and Outlines backends through Jinja2 templates that incorporate type descriptions, the mask placeholder, and per-entity confidence scoring.

Both bridges implement integrate (storing results in doc.results and optionally overwriting doc.text) and consolidate (merging chunk-level outputs into document-level results).

Implementing PII Masking in Your Pipeline

Basic Usage with Any Model Backend

To mask PII in documents, instantiate the PIIMasking task with any supported model wrapper and add it to your pipeline:

from sieves import Doc, Pipeline, tasks

# Prepare documents

docs = [
    Doc(text="John Doe's email is john.doe@example.com."),
    Doc(text="My SSN is 123-45-6789.")
]

# Initialize task with your model wrapper (DSPy, LangChain, or Outlines)

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    batch_size=-1,  # Process all documents at once

)

pipe = Pipeline(task)
masked_docs = list(pipe(docs))

# Access results

for doc in masked_docs:
    result = doc.results["PIIMasking"]
    print("Masked:", result.masked_text)
    for entity in result.pii_entities:
        print(f" • {entity.entity_type}: {entity.text} (score={entity.score})")

Configuring Custom PII Types

You can restrict or define custom entity types by passing a dictionary mapping type names to descriptions. When using a dictionary, the bridge automatically injects an XML-style <pii_type_descriptions> block into the prompt to guide the model:

custom_pii = {
    "EMAIL": "Email addresses",
    "PHONE": "Phone numbers",
    "SSN": "Social security numbers",
    "CREDIT_CARD": "Credit card numbers",
}

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    pii_types=custom_pii,  # Dictionary format adds descriptions to prompts

)

pipe = Pipeline(task)
masked = list(pipe(docs))

Using Few-Shot Prompting

Improve detection accuracy by providing exemplar documents through FewshotExample objects. The framework automatically flattens these into the prompt template:

from sieves.tasks.predictive import pii_masking

fewshot_examples = [
    pii_masking.FewshotExample(
        text="Contact Jane at jane@example.org.",
        masked_text="Contact [MASKED] at [MASKED].",
        pii_entities=[
            pii_masking.PIIEntity(entity_type="PERSON", text="Jane", score=1.0),
            pii_masking.PIIEntity(entity_type="EMAIL", text="jane@example.org", score=1.0),
        ],
    )
]

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    fewshot_examples=fewshot_examples,
)

pipe = Pipeline(task)

Advanced Features and Evaluation

Exporting Results to Hugging Face Datasets

Convert your masked documents into a Hugging Face dataset for downstream analysis or training:

task = tasks.predictive.PIIMasking(model=model_wrapper)
pipe = Pipeline(task)
processed_docs = list(pipe(my_documents))

# Export to HF Dataset with text and masked_text columns

hf_dataset = task.to_hf_dataset(processed_docs)
print(hf_dataset[0])  # {'text': '...', 'masked_text': '...'}

Evaluating Masking Accuracy

The task implements corpus-level micro-F1 evaluation based on exact (entity_type, text) matches. To evaluate against ground truth, attach gold-standard results to documents before calling evaluate:

from sieves import Doc
from sieves.tasks.predictive.schemas.pii_masking import Result, PIIEntity

# Document with ground truth

doc = Doc(text="My email is alice@example.com")
gold_result = Result(
    masked_text="My email is [MASKED]",
    pii_entities=[PIIEntity(entity_type="EMAIL", text="alice@example.com")],
)
doc.gold["pii"] = gold_result

# After running pipeline, model prediction exists in doc.results["pii"]

task = tasks.predictive.PIIMasking(model=model_wrapper, task_id="pii")
report = task.evaluate([doc])
print("F1 Score:", report.metrics[task.metric])  # 1.0 for perfect match

Serialization and State Management

The PIIMasking task supports full serialization through its _state property, which preserves the pii_types configuration (including dictionary descriptions) and ModelSettings. This enables pipeline checkpointing and restoration without losing custom entity type definitions or few-shot example configurations. Override inference modes via ModelSettings(inference_mode=...) to force specific wrapper behaviors such as outlines.InferenceMode.json.

Summary

Frequently Asked Questions

What model backends are supported for PII masking in Sieves?

Sieves supports three model wrappers for PII masking: DSPy, LangChain, and Outlines. Each backend uses a specialized bridge class—DSPyPIIMasking for DSPy implementations and PydanticPIIMasking for both LangChain and Outlines—to handle prompt construction and result parsing while maintaining identical output schemas across all providers.

How does the masking placeholder work when anonymizing text?

The bridge layer automatically prepares a [MASKED] placeholder that replaces detected PII spans in the output text. When the overwrite parameter is set to True in the PIIMasking constructor, the bridge updates doc.text with the masked version during the integrate phase; otherwise, the original text remains unchanged while the masked version is stored in doc.results["PIIMasking"].masked_text.

Can I customize which PII types the model should detect?

Yes, you can specify custom entity types by passing the pii_types parameter as either a list of strings or a dictionary mapping type names to descriptions. When using a dictionary, the bridge injects an XML-style <pii_type_descriptions> block into the Jinja2 or DSPy prompt template, providing the LLM with human-readable context for each entity category you want to detect.

How is PII masking performance measured in the evaluation framework?

The task implements _compute_metrics to calculate corpus-level micro-F1 scores based on exact matches of (entity_type, text) tuples between predicted and gold-standard entities. This metric evaluates both precision and recall at the entity level, requiring perfect string matches for both the entity type classification and the extracted text span to count as a true positive.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →