How to Handle PII Masking and Anonymization in Sieves: A Complete Developer Guide
Sieves provides a dedicated PIIMasking task that automatically detects, masks, and scores personally identifiable information using pluggable model backends including DSPy, LangChain, and Outlines.
The mantisai/sieves library offers a production-ready solution for PII masking and anonymization through its predictive task framework. The PIIMasking class integrates seamlessly into the Pipeline architecture, providing consistent entity detection and text redaction across multiple LLM providers while maintaining strict separation between data models, orchestration logic, and model-specific implementations.
Architecture of the PII Masking System
The PII masking implementation follows a three-layer architecture that separates concerns between data validation, task orchestration, and model interaction.
Data Models and Schemas
At the foundation lies the schema layer in sieves/tasks/predictive/schemas/pii_masking.py, which defines three critical Pydantic models:
PIIEntity– A frozen model describing a single PII span withentity_type,text, and an optionalscoreattribute for confidence levels.Result– The unified task output containing themasked_textversion and a list ofPIIEntityobjects representing detected entities.FewshotExample– A wrapper for few-shot prompting that exposestext,masked_text, andpii_entitiesfields.
These schemas enforce type safety across all model-wrapper bridges, ensuring that DSPy, LangChain, and Outlines implementations return identically structured results regardless of underlying LLM differences.
Task Orchestration Layer
The core task logic resides in sieves/tasks/predictive/pii_masking/core.py, where the PIIMasking class subclasses PredictiveTask to provide:
- Flexible configuration via constructor parameters including
model,pii_types(accepting lists, dictionaries, orNonefor defaults),overwriteto replace original text, customprompt_instructions,fewshot_examples, andModelSettingsfor inference mode overrides. - Prompt signature defined by returning the
Resultmodel fromprompt_signature, standardizing the expected output structure. - Corpus-level evaluation through
_compute_metrics, which calculates micro-F1 scores based on exact matches of(entity_type, text)tuples, with a specialized_evaluate_dspy_examplefor DSPy-specific per-example evaluation. - Dataset export via
to_hf_dataset, creating a Hugging Facedatasets.Datasetwithtextandmasked_textcolumns using the version string fromConfig. - Serialization support through
_state, which preserves thepii_typesconfiguration (including dictionary descriptions when provided) for pipeline persistence.
The core intentionally contains no model-specific logic, delegating all LLM interactions to the bridge layer via _init_bridge.
Bridge Abstraction for Model Backends
The bridge layer in sieves/tasks/predictive/pii_masking/bridges.py abstracts model-specific implementations behind the PIIMaskingBridge base class:
- Placeholder management – The base bridge prepares the
[MASKED]placeholder used to replace sensitive text spans. - PII type parsing – Handles
pii_typesas either a list or dictionary, with dictionary entries generating human-readable descriptions injected into prompts via an XML-style<pii_type_descriptions>block. - Runtime type generation – Creates a runtime
PIIEntityclass with a literal union of allowed entity types for strict validation.
Concrete implementations include:
DSPyPIIMasking– Constructs DSPyPromptSignature/Resultpairs, extracts entities viares.pii_entities, and consolidates chunked results usingMultiEntityConsolidation.PydanticPIIMasking– Powers both LangChain and Outlines backends through Jinja2 templates that incorporate type descriptions, the mask placeholder, and per-entity confidence scoring.
Both bridges implement integrate (storing results in doc.results and optionally overwriting doc.text) and consolidate (merging chunk-level outputs into document-level results).
Implementing PII Masking in Your Pipeline
Basic Usage with Any Model Backend
To mask PII in documents, instantiate the PIIMasking task with any supported model wrapper and add it to your pipeline:
from sieves import Doc, Pipeline, tasks
# Prepare documents
docs = [
Doc(text="John Doe's email is john.doe@example.com."),
Doc(text="My SSN is 123-45-6789.")
]
# Initialize task with your model wrapper (DSPy, LangChain, or Outlines)
task = tasks.predictive.PIIMasking(
model=model_wrapper,
batch_size=-1, # Process all documents at once
)
pipe = Pipeline(task)
masked_docs = list(pipe(docs))
# Access results
for doc in masked_docs:
result = doc.results["PIIMasking"]
print("Masked:", result.masked_text)
for entity in result.pii_entities:
print(f" • {entity.entity_type}: {entity.text} (score={entity.score})")
Configuring Custom PII Types
You can restrict or define custom entity types by passing a dictionary mapping type names to descriptions. When using a dictionary, the bridge automatically injects an XML-style <pii_type_descriptions> block into the prompt to guide the model:
custom_pii = {
"EMAIL": "Email addresses",
"PHONE": "Phone numbers",
"SSN": "Social security numbers",
"CREDIT_CARD": "Credit card numbers",
}
task = tasks.predictive.PIIMasking(
model=model_wrapper,
pii_types=custom_pii, # Dictionary format adds descriptions to prompts
)
pipe = Pipeline(task)
masked = list(pipe(docs))
Using Few-Shot Prompting
Improve detection accuracy by providing exemplar documents through FewshotExample objects. The framework automatically flattens these into the prompt template:
from sieves.tasks.predictive import pii_masking
fewshot_examples = [
pii_masking.FewshotExample(
text="Contact Jane at jane@example.org.",
masked_text="Contact [MASKED] at [MASKED].",
pii_entities=[
pii_masking.PIIEntity(entity_type="PERSON", text="Jane", score=1.0),
pii_masking.PIIEntity(entity_type="EMAIL", text="jane@example.org", score=1.0),
],
)
]
task = tasks.predictive.PIIMasking(
model=model_wrapper,
fewshot_examples=fewshot_examples,
)
pipe = Pipeline(task)
Advanced Features and Evaluation
Exporting Results to Hugging Face Datasets
Convert your masked documents into a Hugging Face dataset for downstream analysis or training:
task = tasks.predictive.PIIMasking(model=model_wrapper)
pipe = Pipeline(task)
processed_docs = list(pipe(my_documents))
# Export to HF Dataset with text and masked_text columns
hf_dataset = task.to_hf_dataset(processed_docs)
print(hf_dataset[0]) # {'text': '...', 'masked_text': '...'}
Evaluating Masking Accuracy
The task implements corpus-level micro-F1 evaluation based on exact (entity_type, text) matches. To evaluate against ground truth, attach gold-standard results to documents before calling evaluate:
from sieves import Doc
from sieves.tasks.predictive.schemas.pii_masking import Result, PIIEntity
# Document with ground truth
doc = Doc(text="My email is alice@example.com")
gold_result = Result(
masked_text="My email is [MASKED]",
pii_entities=[PIIEntity(entity_type="EMAIL", text="alice@example.com")],
)
doc.gold["pii"] = gold_result
# After running pipeline, model prediction exists in doc.results["pii"]
task = tasks.predictive.PIIMasking(model=model_wrapper, task_id="pii")
report = task.evaluate([doc])
print("F1 Score:", report.metrics[task.metric]) # 1.0 for perfect match
Serialization and State Management
The PIIMasking task supports full serialization through its _state property, which preserves the pii_types configuration (including dictionary descriptions) and ModelSettings. This enables pipeline checkpointing and restoration without losing custom entity type definitions or few-shot example configurations. Override inference modes via ModelSettings(inference_mode=...) to force specific wrapper behaviors such as outlines.InferenceMode.json.
Summary
- The
PIIMaskingtask insieves/tasks/predictive/pii_masking/core.pyprovides a unified interface for PII detection and anonymization across DSPy, LangChain, and Outlines backends. - Data models in
sieves/tasks/predictive/schemas/pii_masking.pyenforce consistency through frozen Pydantic schemas for entities, results, and few-shot examples. - Bridge classes in
sieves/tasks/predictive/pii_masking/bridges.pyhandle model-specific prompting, the[MASKED]placeholder substitution, and result consolidation while supporting custom PII type dictionaries with XML-style description injection. - The task calculates micro-F1 metrics based on exact entity matches and supports export to Hugging Face datasets via
to_hf_dataset. - Comprehensive test coverage in
sieves/tests/tasks/predictive/test_pii_masking.pyverifies serialization, evaluation, inference-mode overrides, and integration with all supported model types.
Frequently Asked Questions
What model backends are supported for PII masking in Sieves?
Sieves supports three model wrappers for PII masking: DSPy, LangChain, and Outlines. Each backend uses a specialized bridge class—DSPyPIIMasking for DSPy implementations and PydanticPIIMasking for both LangChain and Outlines—to handle prompt construction and result parsing while maintaining identical output schemas across all providers.
How does the masking placeholder work when anonymizing text?
The bridge layer automatically prepares a [MASKED] placeholder that replaces detected PII spans in the output text. When the overwrite parameter is set to True in the PIIMasking constructor, the bridge updates doc.text with the masked version during the integrate phase; otherwise, the original text remains unchanged while the masked version is stored in doc.results["PIIMasking"].masked_text.
Can I customize which PII types the model should detect?
Yes, you can specify custom entity types by passing the pii_types parameter as either a list of strings or a dictionary mapping type names to descriptions. When using a dictionary, the bridge injects an XML-style <pii_type_descriptions> block into the Jinja2 or DSPy prompt template, providing the LLM with human-readable context for each entity category you want to detect.
How is PII masking performance measured in the evaluation framework?
The task implements _compute_metrics to calculate corpus-level micro-F1 scores based on exact matches of (entity_type, text) tuples between predicted and gold-standard entities. This metric evaluates both precision and recall at the entity level, requiring perfect string matches for both the entity type classification and the extracted text span to count as a true positive.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →