# How to Tune Batch Size and Optimize Performance in Heretic

> Learn how to tune batch size and optimize performance in Heretic. automatically benchmark for optimal throughput or set a max batch size to prevent memory errors. Boost your computational efficiency.

- Repository: [Philipp Emanuel Weidmann/heretic](https://github.com/p-e-w/heretic)
- Tags: performance
- Published: 2026-02-19

---

**To tune batch size and optimize performance in Heretic, set `batch_size = 0` to enable automatic benchmarking that selects the optimal throughput, or manually specify a fixed value and cap it with `max_batch_size` to prevent out-of-memory errors.**

Heretic processes prompts in parallel batches, and the size of each batch directly impacts GPU memory usage and throughput. This guide explains how to configure batch size settings in the `p-e-w/heretic` repository to maximize performance on your hardware.

## Where Batch Size Configuration Lives

Heretic exposes batch size controls through a layered configuration system that combines default values, user overrides, and runtime detection.

### Configuration Files

Default values are defined in [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml) at lines 30-34, where `batch_size = 0` enables automatic tuning and `max_batch_size` sets an upper limit:

```toml
batch_size = 0          # 0 = auto-detect optimal batch size

max_batch_size = 128    # hard cap for auto-tuning search

```

You can override these by creating a [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) file in your working directory or passing command-line arguments.

### Settings Model

The `Settings` Pydantic model in [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) (lines 118-125) validates and exposes these fields:

```python
class Settings(BaseModel):
    batch_size: int = 0
    max_batch_size: int = 128
    # ... other settings

```

This model ensures that `batch_size` and `max_batch_size` are properly typed and accessible throughout the application.

## Automatic Batch Size Tuning

When `batch_size` is set to `0`, Heretic runs a built-in benchmark to find the optimal throughput before processing your actual workload.

### The Benchmarking Process

The auto-tuning logic resides in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) at lines 332-376. The process follows these steps:

1. **Warm-up** – The first call to `model.get_responses` builds the computation graph without timing.
2. **Benchmark** – Heretic repeatedly doubles the batch size, measures tokens-per-second on a small set of *good* prompts, and records performance.
3. **Selection** – The batch size yielding the highest tokens-per-second is stored as `best_batch_size`.
4. **Application** – The chosen size is written back to `settings.batch_size` and used for all subsequent batched calls.

The benchmark automatically adapts to your hardware because it uses actual model generation, including tokenizer encoding and GPU kernel execution. If a batch size causes an out-of-memory (OOM) error, the loop terminates and the last successful size is selected.

### Batch Processing Implementation

Once the size is determined, the `batchify` helper in [`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py) (lines 33-34) splits prompt lists into sub-lists of the target length:

```python
def batchify(prompts: list[Prompt], batch_size: int) -> list[list[Prompt]]:
    return [prompts[i:i + batch_size] for i in range(0, len(prompts), batch_size)]

```

Higher-level methods like `Model.get_responses_batched` in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) (lines 94-101) iterate over these batches to perform parallel inference.

## Manual Batch Size Configuration

For reproducible behavior or specific hardware constraints, you can disable auto-tuning and set fixed values.

### Forcing a Specific Batch Size

To skip the benchmark entirely, set `batch_size` to a non-zero integer in your [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml):

```toml
batch_size = 16
max_batch_size = 128  # ignored when batch_size != 0, but kept for reference

```

This configuration forces Heretic to process exactly 16 prompts per batch, regardless of throughput benchmarks.

### Capping the Auto-Tuning Search

If you experience OOM errors during the benchmark phase, constrain the search space by lowering `max_batch_size`:

```toml
batch_size = 0
max_batch_size = 32

```

Now the auto-tuner will never attempt a batch larger than 32, preventing crashes on memory-constrained GPUs like the RTX 3060 12GB or T4 16GB.

## Performance Optimization Strategies

Beyond batch size, several configuration options interact with memory usage and throughput.

### Combining Quantization with Larger Batches

Enabling 4-bit quantization via `bitsandbytes` reduces model memory footprint, often allowing you to increase batch size without OOM errors:

```toml
quantization = "bnb_4bit"
batch_size = 0
max_batch_size = 256

```

With the model weights compressed, the auto-tuner can explore batches up to 256, potentially achieving higher tokens-per-second through better GPU utilization.

### Adjusting Response Length

The `max_response_length` parameter affects how long each generation runs. Longer sequences increase per-prompt compute time, which may reduce the optimal batch size for your specific workload. If you increase `max_response_length`, consider re-running the auto-tuner to find the new sweet spot.

### Memory-Related Settings

Other settings that influence available VRAM for batch processing include:

- **`winsorization_quantile`** – Affects statistical processing overhead
- **`orthogonalize_direction`** – Impacts memory during direction calculations
- **`row_normalization`** and **`full_normalization_lora_rank`** – Influence LoRA adapter memory footprints

Modifying these can indirectly affect how much memory remains available for batch processing.

## Code Examples

### Use Default Auto-Tuning

Run Heretic with no configuration changes to enable automatic batch size detection:

```bash
heretic meta-llama/Llama-2-7b-chat-hf

```

Heretic reads `batch_size = 0` from [`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml), executes the benchmark in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py), and selects the optimal throughput.

### Force a Fixed Batch Size

Create [`config.toml`](https://github.com/p-e-w/heretic/blob/main/config.toml) to disable auto-tuning and use exactly 16 prompts per batch:

```toml
batch_size = 16
max_batch_size = 128

```

Execute with your custom configuration:

```bash
heretic meta-llama/Llama-2-7b-chat-hf --config config.toml

```

### Prevent OOM on Memory-Constrained GPUs

Limit the auto-tuner to batches of 32 or smaller for 12-16GB VRAM cards:

```toml
batch_size = 0
max_batch_size = 32

```

This ensures the benchmark loop in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) terminates before attempting batches that would exhaust memory.

### Maximize Throughput with Quantization

Combine 4-bit quantization with an increased batch size cap to leverage reduced memory footprint:

```toml
quantization = "bnb_4bit"
batch_size = 0
max_batch_size = 256

```

With model weights compressed via `bitsandbytes`, the auto-tuner can explore batches up to 256, often achieving higher tokens-per-second on high-VRAM GPUs like the A100 or RTX 4090.

## Key Source Files

Understanding these files helps you customize batch behavior beyond configuration files:

- **[`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py)** – Defines the `Settings` Pydantic model that validates `batch_size` and `max_batch_size` from TOML or CLI arguments.
- **[`config.default.toml`](https://github.com/p-e-w/heretic/blob/main/config.default.toml)** – Contains default values including `batch_size = 0` and `max_batch_size = 128`.
- **[`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py)** – Implements the auto-tuning benchmark loop (lines 332-376) that determines optimal batch size through empirical measurement.
- **[`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py)** – Contains the `batchify` helper function that splits prompt lists into batch-sized chunks.
- **[`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py)** – Houses `get_responses_batched` and related methods that execute inference across batches created by `batchify`.

## Summary

- **Auto-tuning is default**: Set `batch_size = 0` in your TOML configuration to enable empirical benchmarking that automatically selects the highest throughput batch size for your GPU.
- **Cap with `max_batch_size`**: Constrain the auto-tuner on memory-limited hardware by setting `max_batch_size` to a safe value (e.g., 32 for 12GB cards).
- **Manual override**: Set `batch_size` to a non-zero integer to skip benchmarking and use a fixed batch size for reproducible behavior.
- **Combine with quantization**: Enable `quantization = "bnb_4bit"` to reduce model memory footprint, allowing the auto-tuner to explore larger batches (up to 256 or higher) for increased throughput.
- **Core implementation**: The tuning logic resides in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py), while batch splitting is handled by `batchify` in [`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py) and consumed by `get_responses_batched` in [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py).

## Frequently Asked Questions

### How does Heretic automatically determine the optimal batch size?

When `batch_size` is set to `0`, Heretic executes a benchmark loop in [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) (lines 332-376) before processing your actual workload. It repeatedly doubles the batch size, measures tokens-per-second on a small set of reference prompts, and selects the size that yields the highest throughput, automatically stopping if it encounters GPU memory limits.

### What is the difference between `batch_size` and `max_batch_size` in Heretic?

The `batch_size` parameter controls how many prompts are processed simultaneously; setting it to `0` enables auto-tuning, while a non-zero value forces a fixed batch size. The `max_batch_size` parameter acts as a safety cap that limits how large the auto-tuner can grow the batch during its benchmark phase, preventing out-of-memory errors on GPUs with limited VRAM.

### Can I use a large batch size with a large language model on limited VRAM?

Yes, but you must reduce memory pressure through quantization or lower the batch size cap. Enabling `quantization = "bnb_4bit"` in your configuration significantly reduces model weight memory, often allowing you to increase `max_batch_size` to 256 or higher even on consumer GPUs. Without quantization, you should set `max_batch_size` to conservative values like 16 or 32 for 12-16GB cards.

### Where is the batch processing logic implemented in the Heretic source code?

The batch processing pipeline spans four key files: [`src/heretic/main.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/main.py) contains the auto-tuning benchmark logic; [`src/heretic/utils.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/utils.py) defines the `batchify` helper that splits prompt lists into batches; [`src/heretic/model.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/model.py) implements `get_responses_batched` which executes inference across batches; and [`src/heretic/config.py`](https://github.com/p-e-w/heretic/blob/main/src/heretic/config.py) validates the `batch_size` and `max_batch_size` settings through the Pydantic `Settings` model.