# How to Handle PII Masking and Anonymization in Sieves: A Complete Developer Guide

> Learn to handle PII masking and anonymization in Sieves. This guide explores the PIIMasking task for automatic PII detection, masking, and scoring with DSPy LangChain and Outlines.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Sieves provides a dedicated PIIMasking task that automatically detects, masks, and scores personally identifiable information using pluggable model backends including DSPy, LangChain, and Outlines.**

The mantisai/sieves library offers a production-ready solution for PII masking and anonymization through its predictive task framework. The `PIIMasking` class integrates seamlessly into the `Pipeline` architecture, providing consistent entity detection and text redaction across multiple LLM providers while maintaining strict separation between data models, orchestration logic, and model-specific implementations.

## Architecture of the PII Masking System

The PII masking implementation follows a three-layer architecture that separates concerns between data validation, task orchestration, and model interaction.

### Data Models and Schemas

At the foundation lies the schema layer in [`sieves/tasks/predictive/schemas/pii_masking.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/schemas/pii_masking.py), which defines three critical Pydantic models:

- **`PIIEntity`** – A frozen model describing a single PII span with `entity_type`, `text`, and an optional `score` attribute for confidence levels.
- **`Result`** – The unified task output containing the `masked_text` version and a list of `PIIEntity` objects representing detected entities.
- **`FewshotExample`** – A wrapper for few-shot prompting that exposes `text`, `masked_text`, and `pii_entities` fields.

These schemas enforce type safety across all model-wrapper bridges, ensuring that DSPy, LangChain, and Outlines implementations return identically structured results regardless of underlying LLM differences.

### Task Orchestration Layer

The core task logic resides in [`sieves/tasks/predictive/pii_masking/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/pii_masking/core.py), where the `PIIMasking` class subclasses `PredictiveTask` to provide:

- **Flexible configuration** via constructor parameters including `model`, `pii_types` (accepting lists, dictionaries, or `None` for defaults), `overwrite` to replace original text, custom `prompt_instructions`, `fewshot_examples`, and `ModelSettings` for inference mode overrides.
- **Prompt signature** defined by returning the `Result` model from `prompt_signature`, standardizing the expected output structure.
- **Corpus-level evaluation** through `_compute_metrics`, which calculates micro-F1 scores based on exact matches of `(entity_type, text)` tuples, with a specialized `_evaluate_dspy_example` for DSPy-specific per-example evaluation.
- **Dataset export** via `to_hf_dataset`, creating a Hugging Face `datasets.Dataset` with `text` and `masked_text` columns using the version string from `Config`.
- **Serialization support** through `_state`, which preserves the `pii_types` configuration (including dictionary descriptions when provided) for pipeline persistence.

The core intentionally contains no model-specific logic, delegating all LLM interactions to the bridge layer via `_init_bridge`.

### Bridge Abstraction for Model Backends

The bridge layer in [`sieves/tasks/predictive/pii_masking/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/pii_masking/bridges.py) abstracts model-specific implementations behind the `PIIMaskingBridge` base class:

- **Placeholder management** – The base bridge prepares the `[MASKED]` placeholder used to replace sensitive text spans.
- **PII type parsing** – Handles `pii_types` as either a list or dictionary, with dictionary entries generating human-readable descriptions injected into prompts via an XML-style `<pii_type_descriptions>` block.
- **Runtime type generation** – Creates a runtime `PIIEntity` class with a literal union of allowed entity types for strict validation.

Concrete implementations include:

- **`DSPyPIIMasking`** – Constructs DSPy `PromptSignature`/`Result` pairs, extracts entities via `res.pii_entities`, and consolidates chunked results using `MultiEntityConsolidation`.
- **`PydanticPIIMasking`** – Powers both LangChain and Outlines backends through Jinja2 templates that incorporate type descriptions, the mask placeholder, and per-entity confidence scoring.

Both bridges implement `integrate` (storing results in `doc.results` and optionally overwriting `doc.text`) and `consolidate` (merging chunk-level outputs into document-level results).

## Implementing PII Masking in Your Pipeline

### Basic Usage with Any Model Backend

To mask PII in documents, instantiate the `PIIMasking` task with any supported model wrapper and add it to your pipeline:

```python
from sieves import Doc, Pipeline, tasks

# Prepare documents

docs = [
    Doc(text="John Doe's email is john.doe@example.com."),
    Doc(text="My SSN is 123-45-6789.")
]

# Initialize task with your model wrapper (DSPy, LangChain, or Outlines)

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    batch_size=-1,  # Process all documents at once

)

pipe = Pipeline(task)
masked_docs = list(pipe(docs))

# Access results

for doc in masked_docs:
    result = doc.results["PIIMasking"]
    print("Masked:", result.masked_text)
    for entity in result.pii_entities:
        print(f" • {entity.entity_type}: {entity.text} (score={entity.score})")

```

### Configuring Custom PII Types

You can restrict or define custom entity types by passing a dictionary mapping type names to descriptions. When using a dictionary, the bridge automatically injects an XML-style `<pii_type_descriptions>` block into the prompt to guide the model:

```python
custom_pii = {
    "EMAIL": "Email addresses",
    "PHONE": "Phone numbers",
    "SSN": "Social security numbers",
    "CREDIT_CARD": "Credit card numbers",
}

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    pii_types=custom_pii,  # Dictionary format adds descriptions to prompts

)

pipe = Pipeline(task)
masked = list(pipe(docs))

```

### Using Few-Shot Prompting

Improve detection accuracy by providing exemplar documents through `FewshotExample` objects. The framework automatically flattens these into the prompt template:

```python
from sieves.tasks.predictive import pii_masking

fewshot_examples = [
    pii_masking.FewshotExample(
        text="Contact Jane at jane@example.org.",
        masked_text="Contact [MASKED] at [MASKED].",
        pii_entities=[
            pii_masking.PIIEntity(entity_type="PERSON", text="Jane", score=1.0),
            pii_masking.PIIEntity(entity_type="EMAIL", text="jane@example.org", score=1.0),
        ],
    )
]

task = tasks.predictive.PIIMasking(
    model=model_wrapper,
    fewshot_examples=fewshot_examples,
)

pipe = Pipeline(task)

```

## Advanced Features and Evaluation

### Exporting Results to Hugging Face Datasets

Convert your masked documents into a Hugging Face dataset for downstream analysis or training:

```python
task = tasks.predictive.PIIMasking(model=model_wrapper)
pipe = Pipeline(task)
processed_docs = list(pipe(my_documents))

# Export to HF Dataset with text and masked_text columns

hf_dataset = task.to_hf_dataset(processed_docs)
print(hf_dataset[0])  # {'text': '...', 'masked_text': '...'}

```

### Evaluating Masking Accuracy

The task implements corpus-level micro-F1 evaluation based on exact `(entity_type, text)` matches. To evaluate against ground truth, attach gold-standard results to documents before calling `evaluate`:

```python
from sieves import Doc
from sieves.tasks.predictive.schemas.pii_masking import Result, PIIEntity

# Document with ground truth

doc = Doc(text="My email is alice@example.com")
gold_result = Result(
    masked_text="My email is [MASKED]",
    pii_entities=[PIIEntity(entity_type="EMAIL", text="alice@example.com")],
)
doc.gold["pii"] = gold_result

# After running pipeline, model prediction exists in doc.results["pii"]

task = tasks.predictive.PIIMasking(model=model_wrapper, task_id="pii")
report = task.evaluate([doc])
print("F1 Score:", report.metrics[task.metric])  # 1.0 for perfect match

```

### Serialization and State Management

The `PIIMasking` task supports full serialization through its `_state` property, which preserves the `pii_types` configuration (including dictionary descriptions) and `ModelSettings`. This enables pipeline checkpointing and restoration without losing custom entity type definitions or few-shot example configurations. Override inference modes via `ModelSettings(inference_mode=...)` to force specific wrapper behaviors such as [`outlines.InferenceMode.json`](https://github.com/mantisai/sieves/blob/main/outlines.InferenceMode.json).

## Summary

- The `PIIMasking` task in [`sieves/tasks/predictive/pii_masking/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/pii_masking/core.py) provides a unified interface for PII detection and anonymization across DSPy, LangChain, and Outlines backends.
- **Data models** in [`sieves/tasks/predictive/schemas/pii_masking.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/schemas/pii_masking.py) enforce consistency through frozen Pydantic schemas for entities, results, and few-shot examples.
- **Bridge classes** in [`sieves/tasks/predictive/pii_masking/bridges.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/pii_masking/bridges.py) handle model-specific prompting, the `[MASKED]` placeholder substitution, and result consolidation while supporting custom PII type dictionaries with XML-style description injection.
- The task calculates **micro-F1 metrics** based on exact entity matches and supports export to Hugging Face datasets via `to_hf_dataset`.
- Comprehensive test coverage in [`sieves/tests/tasks/predictive/test_pii_masking.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/tasks/predictive/test_pii_masking.py) verifies serialization, evaluation, inference-mode overrides, and integration with all supported model types.

## Frequently Asked Questions

### What model backends are supported for PII masking in Sieves?

Sieves supports three model wrappers for PII masking: **DSPy**, **LangChain**, and **Outlines**. Each backend uses a specialized bridge class—`DSPyPIIMasking` for DSPy implementations and `PydanticPIIMasking` for both LangChain and Outlines—to handle prompt construction and result parsing while maintaining identical output schemas across all providers.

### How does the masking placeholder work when anonymizing text?

The bridge layer automatically prepares a `[MASKED]` placeholder that replaces detected PII spans in the output text. When the `overwrite` parameter is set to `True` in the `PIIMasking` constructor, the bridge updates `doc.text` with the masked version during the `integrate` phase; otherwise, the original text remains unchanged while the masked version is stored in `doc.results["PIIMasking"].masked_text`.

### Can I customize which PII types the model should detect?

Yes, you can specify custom entity types by passing the `pii_types` parameter as either a list of strings or a dictionary mapping type names to descriptions. When using a dictionary, the bridge injects an XML-style `<pii_type_descriptions>` block into the Jinja2 or DSPy prompt template, providing the LLM with human-readable context for each entity category you want to detect.

### How is PII masking performance measured in the evaluation framework?

The task implements `_compute_metrics` to calculate corpus-level **micro-F1 scores** based on exact matches of `(entity_type, text)` tuples between predicted and gold-standard entities. This metric evaluates both precision and recall at the entity level, requiring perfect string matches for both the entity type classification and the extracted text span to count as a true positive.