How to Tune Batch Size and Optimize Performance in Heretic
To tune batch size and optimize performance in Heretic, set batch_size = 0 to enable automatic benchmarking that selects the optimal throughput, or manually specify a fixed value and cap it with max_batch_size to prevent out-of-memory errors.
Heretic processes prompts in parallel batches, and the size of each batch directly impacts GPU memory usage and throughput. This guide explains how to configure batch size settings in the p-e-w/heretic repository to maximize performance on your hardware.
Where Batch Size Configuration Lives
Heretic exposes batch size controls through a layered configuration system that combines default values, user overrides, and runtime detection.
Configuration Files
Default values are defined in config.default.toml at lines 30-34, where batch_size = 0 enables automatic tuning and max_batch_size sets an upper limit:
batch_size = 0 # 0 = auto-detect optimal batch size
max_batch_size = 128 # hard cap for auto-tuning search
You can override these by creating a config.toml file in your working directory or passing command-line arguments.
Settings Model
The Settings Pydantic model in src/heretic/config.py (lines 118-125) validates and exposes these fields:
class Settings(BaseModel):
batch_size: int = 0
max_batch_size: int = 128
# ... other settings
This model ensures that batch_size and max_batch_size are properly typed and accessible throughout the application.
Automatic Batch Size Tuning
When batch_size is set to 0, Heretic runs a built-in benchmark to find the optimal throughput before processing your actual workload.
The Benchmarking Process
The auto-tuning logic resides in src/heretic/main.py at lines 332-376. The process follows these steps:
- Warm-up – The first call to
model.get_responsesbuilds the computation graph without timing. - Benchmark – Heretic repeatedly doubles the batch size, measures tokens-per-second on a small set of good prompts, and records performance.
- Selection – The batch size yielding the highest tokens-per-second is stored as
best_batch_size. - Application – The chosen size is written back to
settings.batch_sizeand used for all subsequent batched calls.
The benchmark automatically adapts to your hardware because it uses actual model generation, including tokenizer encoding and GPU kernel execution. If a batch size causes an out-of-memory (OOM) error, the loop terminates and the last successful size is selected.
Batch Processing Implementation
Once the size is determined, the batchify helper in src/heretic/utils.py (lines 33-34) splits prompt lists into sub-lists of the target length:
def batchify(prompts: list[Prompt], batch_size: int) -> list[list[Prompt]]:
return [prompts[i:i + batch_size] for i in range(0, len(prompts), batch_size)]
Higher-level methods like Model.get_responses_batched in src/heretic/model.py (lines 94-101) iterate over these batches to perform parallel inference.
Manual Batch Size Configuration
For reproducible behavior or specific hardware constraints, you can disable auto-tuning and set fixed values.
Forcing a Specific Batch Size
To skip the benchmark entirely, set batch_size to a non-zero integer in your config.toml:
batch_size = 16
max_batch_size = 128 # ignored when batch_size != 0, but kept for reference
This configuration forces Heretic to process exactly 16 prompts per batch, regardless of throughput benchmarks.
Capping the Auto-Tuning Search
If you experience OOM errors during the benchmark phase, constrain the search space by lowering max_batch_size:
batch_size = 0
max_batch_size = 32
Now the auto-tuner will never attempt a batch larger than 32, preventing crashes on memory-constrained GPUs like the RTX 3060 12GB or T4 16GB.
Performance Optimization Strategies
Beyond batch size, several configuration options interact with memory usage and throughput.
Combining Quantization with Larger Batches
Enabling 4-bit quantization via bitsandbytes reduces model memory footprint, often allowing you to increase batch size without OOM errors:
quantization = "bnb_4bit"
batch_size = 0
max_batch_size = 256
With the model weights compressed, the auto-tuner can explore batches up to 256, potentially achieving higher tokens-per-second through better GPU utilization.
Adjusting Response Length
The max_response_length parameter affects how long each generation runs. Longer sequences increase per-prompt compute time, which may reduce the optimal batch size for your specific workload. If you increase max_response_length, consider re-running the auto-tuner to find the new sweet spot.
Memory-Related Settings
Other settings that influence available VRAM for batch processing include:
winsorization_quantile– Affects statistical processing overheadorthogonalize_direction– Impacts memory during direction calculationsrow_normalizationandfull_normalization_lora_rank– Influence LoRA adapter memory footprints
Modifying these can indirectly affect how much memory remains available for batch processing.
Code Examples
Use Default Auto-Tuning
Run Heretic with no configuration changes to enable automatic batch size detection:
heretic meta-llama/Llama-2-7b-chat-hf
Heretic reads batch_size = 0 from config.default.toml, executes the benchmark in src/heretic/main.py, and selects the optimal throughput.
Force a Fixed Batch Size
Create config.toml to disable auto-tuning and use exactly 16 prompts per batch:
batch_size = 16
max_batch_size = 128
Execute with your custom configuration:
heretic meta-llama/Llama-2-7b-chat-hf --config config.toml
Prevent OOM on Memory-Constrained GPUs
Limit the auto-tuner to batches of 32 or smaller for 12-16GB VRAM cards:
batch_size = 0
max_batch_size = 32
This ensures the benchmark loop in src/heretic/main.py terminates before attempting batches that would exhaust memory.
Maximize Throughput with Quantization
Combine 4-bit quantization with an increased batch size cap to leverage reduced memory footprint:
quantization = "bnb_4bit"
batch_size = 0
max_batch_size = 256
With model weights compressed via bitsandbytes, the auto-tuner can explore batches up to 256, often achieving higher tokens-per-second on high-VRAM GPUs like the A100 or RTX 4090.
Key Source Files
Understanding these files helps you customize batch behavior beyond configuration files:
src/heretic/config.py– Defines theSettingsPydantic model that validatesbatch_sizeandmax_batch_sizefrom TOML or CLI arguments.config.default.toml– Contains default values includingbatch_size = 0andmax_batch_size = 128.src/heretic/main.py– Implements the auto-tuning benchmark loop (lines 332-376) that determines optimal batch size through empirical measurement.src/heretic/utils.py– Contains thebatchifyhelper function that splits prompt lists into batch-sized chunks.src/heretic/model.py– Housesget_responses_batchedand related methods that execute inference across batches created bybatchify.
Summary
- Auto-tuning is default: Set
batch_size = 0in your TOML configuration to enable empirical benchmarking that automatically selects the highest throughput batch size for your GPU. - Cap with
max_batch_size: Constrain the auto-tuner on memory-limited hardware by settingmax_batch_sizeto a safe value (e.g., 32 for 12GB cards). - Manual override: Set
batch_sizeto a non-zero integer to skip benchmarking and use a fixed batch size for reproducible behavior. - Combine with quantization: Enable
quantization = "bnb_4bit"to reduce model memory footprint, allowing the auto-tuner to explore larger batches (up to 256 or higher) for increased throughput. - Core implementation: The tuning logic resides in
src/heretic/main.py, while batch splitting is handled bybatchifyinsrc/heretic/utils.pyand consumed byget_responses_batchedinsrc/heretic/model.py.
Frequently Asked Questions
How does Heretic automatically determine the optimal batch size?
When batch_size is set to 0, Heretic executes a benchmark loop in src/heretic/main.py (lines 332-376) before processing your actual workload. It repeatedly doubles the batch size, measures tokens-per-second on a small set of reference prompts, and selects the size that yields the highest throughput, automatically stopping if it encounters GPU memory limits.
What is the difference between batch_size and max_batch_size in Heretic?
The batch_size parameter controls how many prompts are processed simultaneously; setting it to 0 enables auto-tuning, while a non-zero value forces a fixed batch size. The max_batch_size parameter acts as a safety cap that limits how large the auto-tuner can grow the batch during its benchmark phase, preventing out-of-memory errors on GPUs with limited VRAM.
Can I use a large batch size with a large language model on limited VRAM?
Yes, but you must reduce memory pressure through quantization or lower the batch size cap. Enabling quantization = "bnb_4bit" in your configuration significantly reduces model weight memory, often allowing you to increase max_batch_size to 256 or higher even on consumer GPUs. Without quantization, you should set max_batch_size to conservative values like 16 or 32 for 12-16GB cards.
Where is the batch processing logic implemented in the Heretic source code?
The batch processing pipeline spans four key files: src/heretic/main.py contains the auto-tuning benchmark logic; src/heretic/utils.py defines the batchify helper that splits prompt lists into batches; src/heretic/model.py implements get_responses_batched which executes inference across batches; and src/heretic/config.py validates the batch_size and max_batch_size settings through the Pydantic Settings model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →