# How to Configure ModelSettings for Inference Modes and Batch Processing in Sieves

> Configure ModelSettings in Sieves for JSON or chain-of-thought inference modes and optimize throughput using the batch size parameter. Learn how to streamline your processing.

- Repository: [Mantis/sieves](https://github.com/mantisai/sieves)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Use `ModelSettings` to set structured generation modes like JSON or chain-of-thought, and control throughput with the `batch_size` parameter on any Task.**

Sieves is a modular document processing framework that unifies LLM wrappers under a common interface. To unlock structured outputs and optimize API costs, you configure `ModelSettings` for inference modes while tuning `batch_size` for batch processing. This guide shows you how to combine both capabilities using the actual source implementation.

## Understanding ModelSettings Architecture

### Core Fields in types.py

The `ModelSettings` class in [`sieves/model_wrappers/types.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py) defines five key fields that control model behavior:

```python
class ModelSettings(pydantic.BaseModel):
    """
    Settings for model for structured generation.

    :param init_kwargs: kwargs passed on to initialization of structured generator.
    :param inference_kwargs: kwargs passed on to inference with structured generator.
    :param config_kwargs: DSPy‑specific configuration kwargs (ignored for other wrappers).
    :param strict: if True, raise on unparsable responses.
    :param inference_mode: specifies the inference mode for the model wrapper.
    """
    init_kwargs: dict[str, Any] | None = None
    inference_kwargs: dict[str, Any] | None = None
    config_kwargs: dict[str, Any] | None = None
    strict: bool = True
    inference_mode: Any | None = None

```

*Source:* [[`sieves/model_wrappers/types.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py)](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/types.py) (lines 8-25)

### Wrapper-Specific Inference Modes

Each model wrapper exposes its own `InferenceMode` enum. You must import the correct enum for your chosen backend.

**Outlines** supports four modes defined in [`sieves/model_wrappers/outlines_.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines_.py):

- `InferenceMode.text` – Raw text generation.
- `InferenceMode.choice` – Select from predefined options (ideal for classification).
- `InferenceMode.regex` – Extract text matching a regex pattern.
- [`InferenceMode.json`](https://github.com/mantisai/sieves/blob/main/InferenceMode.json) – Structured JSON adhering to a Pydantic schema.

*Source:* [[`sieves/model_wrappers/outlines_.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines_.py)](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/outlines_.py) (lines 12-24)

**DSPy** offers three reasoning strategies in [`sieves/model_wrappers/dspy_.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py):

- `InferenceMode.predict` – Direct prediction via `dspy.Predict`.
- `InferenceMode.chain_of_thought` – Chain-of-thought prompting.
- `InferenceMode.react` – ReAct style reasoning.

*Source:* [[`sieves/model_wrappers/dspy_.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py)](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py) (lines 12-22)

When a task initializes, the bridge mediating between the task and wrapper looks up `self._model_settings.inference_mode`. If the value is `None`, the bridge falls back to the wrapper's default. Setting the field explicitly overrides the default, which is verified by the test suite in `test_inference_mode_override`.

*Source:* [[`sieves/tests/tasks/predictive/test_classification.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/tasks/predictive/test_classification.py)](https://github.com/mantisai/sieves/blob/main/sieves/tests/tasks/predictive/test_classification.py) (lines 78-92)

## Configuring Batch Processing with batch_size

While `ModelSettings` controls model behavior, throughput is governed by the `batch_size` parameter on the `Task` base class.

### Task Base Class Implementation

The `Task` constructor in [`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) accepts `batch_size` with the following semantics:

- `batch_size > 0` – Process exactly that many documents per model call.
- `batch_size = -1` – Disable batching; process the entire document set in one call (useful for small datasets or debugging).

```python
def __init__(self, task_id: str | None, include_meta: bool,
             batch_size: int, condition: Callable[[Doc], bool] | None = None):
    """
    :param batch_size: Batch size for processing documents.
                       Use -1 to process all documents at once.
    """
    self._batch_size = batch_size

```

*Source:* [[`sieves/tasks/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/core.py) (lines 24-38)

### Inheritance in Predictive Tasks

All predictive tasks (classification, NER, etc.) inherit from `PredictiveTask`, which passes `batch_size` directly to the `Task` parent. For example, the `Classification` task constructor signature includes `batch_size: int = -1` and forwards it via `super().__init__(..., batch_size=batch_size)`.

*Source:* [[`sieves/tasks/predictive/classification/core.py`](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/classification/core.py)](https://github.com/mantisai/sieves/blob/main/sieves/tasks/predictive/classification/core.py) (lines 49-58)

When a `Task` is called, it materializes documents in batches of size `self._batch_size` (or `sys.maxsize` when `batch_size <= 0`). This batching logic is shared by all wrappers because the bridge receives an already-batched iterator of `Doc` objects.

## Practical Configuration Examples

### Structured JSON Output with Batch Size

This example configures an Outlines-backed classifier to emit JSON and processes documents in batches of eight:

```python
from sieves import Doc, Pipeline
from sieves.model_wrappers import ModelType, ModelSettings
from sieves.tasks import Classification

docs = [Doc(text="Stock markets rallied today."), Doc(text="New vaccine shows promise.")]

# Configure ModelSettings for JSON structured generation

settings = ModelSettings(
    inference_mode="json",          # Must match Outlines InferenceMode.json

    init_kwargs={"temperature": 0.0},
    inference_kwargs={"max_new_tokens": 256},
    strict=True
)

# Create task with explicit batch size

clf = Classification(
    task_id="sentiment",
    labels={"bullish": "Positive market sentiment", "bearish": "Negative market sentiment"},
    model=ModelType.outlines,
    model_settings=settings,
    batch_size=8,                   # Process up to 8 docs per API call

    mode="single"
)

pipeline = Pipeline(clf)
results = list(pipeline(docs))

```

*Source:* [[`sieves/tests/docs/test_getting_started.py`](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_getting_started.py)](https://github.com/mantisai/sieves/blob/main/sieves/tests/docs/test_getting_started.py) (lines 100-110)

### DSPy Chain-of-Thought with Full Batch Processing

For DSPy models, import the specific `InferenceMode` enum and disable batching to process the entire dataset at once:

```python
from sieves.model_wrappers.dspy_ import InferenceMode as DSPyMode
from sieves.model_wrappers import ModelSettings, ModelType
from sieves.tasks import Classification

settings = ModelSettings(
    inference_mode=DSPyMode.chain_of_thought,
    config_kwargs={"lm": "gpt-4"}
)

task = Classification(
    labels=["sports", "politics", "tech"],
    model=ModelType.dspy,
    model_settings=settings,
    batch_size=-1,          # Process all documents in a single call

)

```

*Source:* [[`sieves/model_wrappers/dspy_.py`](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py)](https://github.com/mantisai/sieves/blob/main/sieves/model_wrappers/dspy_.py) (lines 12-22)

### Mixed Batch Sizes in a Pipeline

You can combine tasks with different batch sizes in a single pipeline:

```python
from sieves import Pipeline, Doc
from sieves.tasks import Classification, Chunking

pipeline = Pipeline([
    Chunking(task_id="chunker", batch_size=-1),          # chunk the whole collection at once

    Classification(
        task_id="topic_classifier",
        labels=["tech", "health", "finance"],
        model=ModelType.outlines,
        model_settings=ModelSettings(inference_mode="choice"),
        batch_size=8,                                   # send up to 8 chunks per request

    ),
])

docs = [Doc(text="AI research is advancing fast.") for _ in range(20)]
processed = list(pipeline(docs))   # 20 docs → 3 API calls (8+8+4)

```

## Summary

- **ModelSettings** (defined in [`types.py`](https://github.com/mantisai/sieves/blob/main/types.py)) is the canonical configuration object for structured generation, exposing `inference_mode`, `init_kwargs`, `inference_kwargs`, `config_kwargs`, and `strict`.
- **Inference modes** are wrapper‑specific enums (e.g., [`InferenceMode.json`](https://github.com/mantisai/sieves/blob/main/InferenceMode.json) for Outlines, `InferenceMode.chain_of_thought` for DSPy). Passing the correct enum value to `ModelSettings.inference_mode` overrides the wrapper default.
- **Batch processing** is controlled by the `batch_size` parameter on the `Task` base class ([`core.py`](https://github.com/mantisai/sieves/blob/main/core.py)). Values `>0` enable fixed‑size batches; `-1` disables batching and processes the full document set at once.
- Predictive tasks (classification, NER, etc.) inherit this behavior via `PredictiveTask`, ensuring consistent batch semantics across all model wrappers.

## Frequently Asked Questions

### What happens if I set an unsupported inference mode?

The framework validates the `inference_mode` value against the wrapper's `InferenceMode` enum when the task initializes. If the value does not match a supported member, a `KeyError` or `ValueError` is raised immediately, preventing runtime failures during document processing.

### Can I use different batch sizes for different tasks in the same pipeline?

Yes. Each `Task` maintains its own `batch_size` attribute. When you construct a `Pipeline` with multiple tasks, the pipeline executor respects each task's individual batch size. For example, you can set `batch_size=-1` for a fast preprocessing task and `batch_size=8` for a costly LLM inference task.

### How do I enable chain-of-thought reasoning with DSPy?

Import the `InferenceMode` enum from `sieves.model_wrappers.dspy_` and pass `InferenceMode.chain_of_thought` to `ModelSettings.inference_mode`. This instructs the DSPy wrapper to use `dspy.ChainOfThought` instead of the default `dspy.Predict`, enabling multi-step reasoning without changing any other task configuration.

### Is there a performance penalty for setting batch_size to -1?

Setting `batch_size=-1` disables internal batching, meaning all documents are passed to the model wrapper in a single list. For small datasets or local models, this is often faster because it avoids Python loop overhead. However, for remote APIs with rate limits or payload size restrictions, large batches may trigger errors; in those cases, use an explicit `batch_size` (e.g., 8 or 16) to stay within provider limits.