How to Perform Information Extraction with Single vs Multi-Entity Mode in Sieves
Use mode="single" in InformationExtraction to extract exactly one entity per document (result stored in entity, metric: Accuracy) and mode="multi" to extract all occurrences (result stored in entities, metric: F1).
The InformationExtraction task in the Sieves library enables structured data extraction from unstructured text through a unified predictive interface. When configuring this task, you must specify whether to extract exactly one occurrence or all occurrences of a target entity type using the single vs multi-entity mode parameter.
Understanding Single vs Multi-Entity Extraction Modes
Sieves defines extraction behavior through a Literal["multi", "single"] argument passed to the InformationExtraction constructor at line 73 of sieves/tasks/predictive/information_extraction/core.py. This mode selection determines the output schema, result attribute names, and evaluation metrics used during validation.
Single-Entity Mode (mode="single")
Set mode="single" when your document contains exactly one target entity, such as a document title, primary author, or main subject. In this mode, the task returns a single EntityModel instance stored under doc.results[task_id].entity. According to the source code at line 158, this mode automatically configures Accuracy as the evaluation metric, measuring the correctness of the single prediction.
Multi-Entity Mode (mode="multi")
Use mode="multi" for list-type extractions such as all dates, named entities, or key-phrases found in the document. The bridge creates a dynamic Pydantic model using pydantic.create_model(..., __root__=List[EntityModel]) as implemented in lines 124‑138 of the core file. Results are stored as a list under doc.results[task_id].entities, and the evaluation metric defaults to F1 to account for multiple span predictions.
Implementation Details in Sieves Source Code
The mode parameter propagates through three critical components:
-
Task Core (
sieves/tasks/predictive/information_extraction/core.py): Acceptsmodeat line 73 and stores it inself._mode. The post-processing logic at lines 124‑138 (multi) and lines 169‑170 (single) branches based on this flag to pack raw model outputs correctly. -
Bridge Classes (
sieves/tasks/predictive/information_extraction/bridges.py): TheInformationExtractionBridgebase class and its concrete subclasses receive themodeargument at line 68. The bridge uses this to construct appropriate prompts and parse responses, ensuring the model returns either a single object or a list. -
Factory Method: The
_init_bridgefactory propagates the mode attribute to bridge instances, maintaining consistency across different model wrappers.
Code Examples for Single and Multi-Entity Extraction
Defining the Entity Schema
First, define a Pydantic model describing the entity structure:
from pydantic import BaseModel, Field
class Person(BaseModel):
name: str = Field(..., description="Full name of the person")
age: int | None = Field(None, description="Age if mentioned")
score: float | None = Field(
None, description="Model confidence for this extraction"
)
Single-Entity Extraction
Extract exactly one entity when you expect a unique value per document:
from sieves import Pipeline, Doc
from sieves.tasks.predictive.information_extraction import InformationExtraction
# Create task with mode="single"
extract_author = InformationExtraction(
entity_type=Person,
model="gpt-4o-mini",
mode="single",
)
pipe = Pipeline([extract_author])
doc = Doc(text="This article was written by Alice Johnson, aged 34.")
result_doc = pipe([doc])[0]
# Access single entity via .entity
author = result_doc.results[extract_author.task_id].entity
print(author.name, author.age) # → Alice Johnson 34
Multi-Entity Extraction
Extract all matching entities when multiple instances exist:
# Create task with mode="multi"
extract_people = InformationExtraction(
entity_type=Person,
model="gpt-4o-mini",
mode="multi",
)
pipe = Pipeline([extract_people])
doc = Doc(text="Alice (34) and Bob (28) attended the meeting.")
result_doc = pipe([doc])[0]
# Access list of entities via .entities
people = result_doc.results[extract_people.task_id].entities
for p in people:
print(p.name, p.age)
# → Alice 34
# → Bob 28
Integration in Multi-Task Pipelines
Combine extraction with other tasks like sentiment analysis:
from sieves.tasks.predictive.sentiment_classification import SentimentClassification
sentiment = SentimentClassification(model="gpt-4o-mini")
pipeline = Pipeline([sentiment, extract_people])
doc = Doc(text="Positive vibes from Alice and Bob.")
out = pipeline([doc])[0]
print(out.results[sentiment.task_id].label) # → Positive
print([p.name for p in out.results[extract_people.task_id].entities])
# → ["Alice", "Bob"]
Summary
- Single mode (
mode="single") extracts exactly one entity, stores it inentity, and uses Accuracy for evaluation. - Multi mode (
mode="multi") extracts all entities, stores them inentities, and uses F1 for evaluation. - The
InformationExtractionconstructor accepts the mode parameter at line 73 ofcore.pyand propagates it to bridge classes via_init_bridge. - Bridge implementations in
bridges.pyadjust prompt construction and response parsing based on the mode flag received at line 68. - Result attributes differ between modes: access
.entityfor single mode and.entitiesfor multi-mode to retrieve extracted data.
Frequently Asked Questions
What is the difference between single and multi-entity mode in Sieves?
Single-entity mode returns exactly one occurrence of the target type (the most confident prediction), while multi-entity mode returns all occurrences found in the document. Single mode stores results under the entity attribute and uses Accuracy metrics, whereas multi mode stores results under entities as a list and uses F1 scoring to handle multiple spans.
How does the extraction mode affect evaluation metrics?
According to the implementation at line 158 in sieves/tasks/predictive/information_extraction/core.py, single mode automatically sets self._metric = "Accuracy" to measure single prediction correctness, while multi mode sets self._metric = "F1" to evaluate performance across multiple extracted entity spans.
Can I change the mode after creating an InformationExtraction task?
No, the mode is set during task instantiation and stored in self._mode (line 73 of core.py). To switch extraction behavior, you must create a new InformationExtraction instance with the desired mode parameter and replace the task in your pipeline.
Which result attribute should I access for each mode?
Access doc.results[task_id].entity when using mode="single" to retrieve a single EntityModel instance. Access doc.results[task_id].entities when using mode="multi" to retrieve a list[EntityModel]. The task automatically adjusts the result container based on the mode specified at initialization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →