How to Configure ModelSettings for Inference Modes and Batch Processing in Sieves

Use ModelSettings to set structured generation modes like JSON or chain-of-thought, and control throughput with the batch_size parameter on any Task.

Sieves is a modular document processing framework that unifies LLM wrappers under a common interface. To unlock structured outputs and optimize API costs, you configure ModelSettings for inference modes while tuning batch_size for batch processing. This guide shows you how to combine both capabilities using the actual source implementation.

Understanding ModelSettings Architecture

Core Fields in types.py

The ModelSettings class in sieves/model_wrappers/types.py defines five key fields that control model behavior:

class ModelSettings(pydantic.BaseModel):
    """
    Settings for model for structured generation.

    :param init_kwargs: kwargs passed on to initialization of structured generator.
    :param inference_kwargs: kwargs passed on to inference with structured generator.
    :param config_kwargs: DSPy‑specific configuration kwargs (ignored for other wrappers).
    :param strict: if True, raise on unparsable responses.
    :param inference_mode: specifies the inference mode for the model wrapper.
    """
    init_kwargs: dict[str, Any] | None = None
    inference_kwargs: dict[str, Any] | None = None
    config_kwargs: dict[str, Any] | None = None
    strict: bool = True
    inference_mode: Any | None = None

Source: [sieves/model_wrappers/types.py](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py) (lines 8-25)

Wrapper-Specific Inference Modes

Each model wrapper exposes its own InferenceMode enum. You must import the correct enum for your chosen backend.

Outlines supports four modes defined in sieves/model_wrappers/outlines_.py:

  • InferenceMode.text – Raw text generation.
  • InferenceMode.choice – Select from predefined options (ideal for classification).
  • InferenceMode.regex – Extract text matching a regex pattern.
  • InferenceMode.json – Structured JSON adhering to a Pydantic schema.

Source: [sieves/model_wrappers/outlines_.py](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines_.py) (lines 12-24)

DSPy offers three reasoning strategies in sieves/model_wrappers/dspy_.py:

  • InferenceMode.predict – Direct prediction via dspy.Predict.
  • InferenceMode.chain_of_thought – Chain-of-thought prompting.
  • InferenceMode.react – ReAct style reasoning.

Source: [sieves/model_wrappers/dspy_.py](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py) (lines 12-22)

When a task initializes, the bridge mediating between the task and wrapper looks up self._model_settings.inference_mode. If the value is None, the bridge falls back to the wrapper's default. Setting the field explicitly overrides the default, which is verified by the test suite in test_inference_mode_override.

Source: [sieves/tests/tasks/predictive/test_classification.py](https://github.com/mantisai/sieves/blob/main/sieves/tests/tasks/predictive/test_classification.py) (lines 78-92)

Configuring Batch Processing with batch_size

While ModelSettings controls model behavior, throughput is governed by the batch_size parameter on the Task base class.

Task Base Class Implementation

The Task constructor in sieves/tasks/core.py accepts batch_size with the following semantics:

  • batch_size > 0 – Process exactly that many documents per model call.
  • batch_size = -1 – Disable batching; process the entire document set in one call (useful for small datasets or debugging).
def __init__(self, task_id: str | None, include_meta: bool,
             batch_size: int, condition: Callable[[Doc], bool] | None = None):
    """
    :param batch_size: Batch size for processing documents.
                       Use -1 to process all documents at once.
    """
    self._batch_size = batch_size

Source: [sieves/tasks/core.py](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) (lines 24-38)

Inheritance in Predictive Tasks

All predictive tasks (classification, NER, etc.) inherit from PredictiveTask, which passes batch_size directly to the Task parent. For example, the Classification task constructor signature includes batch_size: int = -1 and forwards it via super().__init__(..., batch_size=batch_size).

Source: [sieves/tasks/predictive/classification/core.py](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/classification/core.py) (lines 49-58)

When a Task is called, it materializes documents in batches of size self._batch_size (or sys.maxsize when batch_size <= 0). This batching logic is shared by all wrappers because the bridge receives an already-batched iterator of Doc objects.

Practical Configuration Examples

Structured JSON Output with Batch Size

This example configures an Outlines-backed classifier to emit JSON and processes documents in batches of eight:

from sieves import Doc, Pipeline
from sieves.model_wrappers import ModelType, ModelSettings
from sieves.tasks import Classification

docs = [Doc(text="Stock markets rallied today."), Doc(text="New vaccine shows promise.")]

# Configure ModelSettings for JSON structured generation

settings = ModelSettings(
    inference_mode="json",          # Must match Outlines InferenceMode.json

    init_kwargs={"temperature": 0.0},
    inference_kwargs={"max_new_tokens": 256},
    strict=True
)

# Create task with explicit batch size

clf = Classification(
    task_id="sentiment",
    labels={"bullish": "Positive market sentiment", "bearish": "Negative market sentiment"},
    model=ModelType.outlines,
    model_settings=settings,
    batch_size=8,                   # Process up to 8 docs per API call

    mode="single"
)

pipeline = Pipeline(clf)
results = list(pipeline(docs))

Source: [sieves/tests/docs/test_getting_started.py](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_getting_started.py) (lines 100-110)

DSPy Chain-of-Thought with Full Batch Processing

For DSPy models, import the specific InferenceMode enum and disable batching to process the entire dataset at once:

from sieves.model_wrappers.dspy_ import InferenceMode as DSPyMode
from sieves.model_wrappers import ModelSettings, ModelType
from sieves.tasks import Classification

settings = ModelSettings(
    inference_mode=DSPyMode.chain_of_thought,
    config_kwargs={"lm": "gpt-4"}
)

task = Classification(
    labels=["sports", "politics", "tech"],
    model=ModelType.dspy,
    model_settings=settings,
    batch_size=-1,          # Process all documents in a single call

)

Source: [sieves/model_wrappers/dspy_.py](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py) (lines 12-22)

Mixed Batch Sizes in a Pipeline

You can combine tasks with different batch sizes in a single pipeline:

from sieves import Pipeline, Doc
from sieves.tasks import Classification, Chunking

pipeline = Pipeline([
    Chunking(task_id="chunker", batch_size=-1),          # chunk the whole collection at once

    Classification(
        task_id="topic_classifier",
        labels=["tech", "health", "finance"],
        model=ModelType.outlines,
        model_settings=ModelSettings(inference_mode="choice"),
        batch_size=8,                                   # send up to 8 chunks per request

    ),
])

docs = [Doc(text="AI research is advancing fast.") for _ in range(20)]
processed = list(pipeline(docs))   # 20 docs → 3 API calls (8+8+4)

Summary

  • ModelSettings (defined in types.py) is the canonical configuration object for structured generation, exposing inference_mode, init_kwargs, inference_kwargs, config_kwargs, and strict.
  • Inference modes are wrapper‑specific enums (e.g., InferenceMode.json for Outlines, InferenceMode.chain_of_thought for DSPy). Passing the correct enum value to ModelSettings.inference_mode overrides the wrapper default.
  • Batch processing is controlled by the batch_size parameter on the Task base class (core.py). Values >0 enable fixed‑size batches; -1 disables batching and processes the full document set at once.
  • Predictive tasks (classification, NER, etc.) inherit this behavior via PredictiveTask, ensuring consistent batch semantics across all model wrappers.

Frequently Asked Questions

What happens if I set an unsupported inference mode?

The framework validates the inference_mode value against the wrapper's InferenceMode enum when the task initializes. If the value does not match a supported member, a KeyError or ValueError is raised immediately, preventing runtime failures during document processing.

Can I use different batch sizes for different tasks in the same pipeline?

Yes. Each Task maintains its own batch_size attribute. When you construct a Pipeline with multiple tasks, the pipeline executor respects each task's individual batch size. For example, you can set batch_size=-1 for a fast preprocessing task and batch_size=8 for a costly LLM inference task.

How do I enable chain-of-thought reasoning with DSPy?

Import the InferenceMode enum from sieves.model_wrappers.dspy_ and pass InferenceMode.chain_of_thought to ModelSettings.inference_mode. This instructs the DSPy wrapper to use dspy.ChainOfThought instead of the default dspy.Predict, enabling multi-step reasoning without changing any other task configuration.

Is there a performance penalty for setting batch_size to -1?

Setting batch_size=-1 disables internal batching, meaning all documents are passed to the model wrapper in a single list. For small datasets or local models, this is often faster because it avoids Python loop overhead. However, for remote APIs with rate limits or payload size restrictions, large batches may trigger errors; in those cases, use an explicit batch_size (e.g., 8 or 16) to stay within provider limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →