# How to Configure Threading for Parallel Search in Turbovec Using `RAYON_NUM_THREADS`

> Optimize Turbovec search performance by configuring threading with RAYON_NUM_THREADS. Learn how to set the environment variable for efficient parallel processing and maximize your hardware utilization. Control thread pool size ...

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: how-to-guide
- Published: 2026-08-22

---

**TLDR: Set the `RAYON_NUM_THREADS` environment variable before importing turbovec to control the Rayon thread pool size used for parallel vector search; the value is clamped to a maximum of 4× hardware parallelism, with a fallback cap of 1024.**

Turbovec is a high-performance Rust-based vector search library with Python bindings that leverages the Rayon crate for CPU-bound parallelism. According to the source code in `RyanCodrai/turbovec`, the number of worker threads in the process-local Rayon pool is entirely controlled by the `RAYON_NUM_THREADS` environment variable. This guide explains exactly how the threading configuration works under the hood and provides runnable Python examples for both manual and automatic thread management.

## Where Rayon Threading Is Initialized

Turbovec configures its Rayon pool during the Python binding initialization. The critical logic lives in **[`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs)**, specifically in the `init_rayon_pool` function. When you run `import turbovec`, the following sequence executes:

1. The code reads `RAYON_NUM_THREADS` from the environment and parses it as a `usize`.
2. It clamps the requested value against a maximum detected cap (`rayon_thread_cap`).
3. If the requested count exceeds the cap, a **RuntimeWarning** is emitted through Python's `warnings` module.
4. The actual pool size is returned by the `desired_threads()` function, which falls back to OS-reported parallelism when the environment variable is absent or unparseable.

The core search operation in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) then queries the pool size at runtime, ensuring every parallel search respects the configured thread count.

## The Thread Cap Logic

Turbovec does not allow unlimited threads. The source code in [`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) defines the cap as **four times the detected hardware parallelism** (lines 1906–1918). If the system's hardware parallelism cannot be determined, the cap falls back to **1024** threads.

Here's how the clamping behaves in practice:

- If you request **8 threads** on a 4-core machine, you get 8 threads (assuming the cap of 16 or more is not exceeded).
- If you request **200 threads** on a 32-core machine, the cap of 128 (4 × 32) reduces the pool to 128 threads.
- If you request an absurd value like 20000, the pool is capped at the maximum and a warning informs you of the reduction.

## How Parallel Search Uses the Thread Count

During a search operation, turbovec's core scheduling logic derives the worker thread count from the Rayon pool. In [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) (lines 1330–1336), the variable `n_threads` is passed to the block-parallel search engine. This means:

- Every search invocation uses the same global thread pool.
- The parallelism is applied across data blocks, making heavy retrieval tasks faster on multi-core systems.
- Changes to the environment variable require a new Python process to take effect.

## Setting the Thread Count: Practical Examples

### Example 1: Explicit Thread Configuration

Set `RAYON_NUM_THREADS` before importing turbovec to guarantee a fixed pool size:

```python
import os

# Use 8 worker threads for the Rayon pool (will be clamped to <= cap)

os.environ["RAYON_NUM_THREADS"] = "8"

import turbovec  # init_rayon_pool runs here and reads the env var

# Now run a parallel search; the pool will have exactly 8 threads

vectors = ...          # (n_vectors, dim) float32 ndarray

queries = ...          # (n_queries, dim) float32 ndarray

k = 10                 # retrieve top-k results

results = turbovec.search(vectors, queries, k)

```

### Example 2: Automatic Hardware Detection

If you omit the environment variable, turbovec falls back to the OS-reported parallelism. This is ideal for ensuring portability across machines:

```python
import turbovec

# Turbovec will use the hardware-detected parallelism (×4 cap) automatically

vectors = ...
queries = ...
results = turbovec.search(vectors, queries, k=5)

```

### Example 3: Handling an Over-Large Request

The clamping produces a warning when the requested value exceeds the cap. You can capture and assert on that warning:

```python
import os, warnings

# Request an unrealistically high thread count

os.environ["RAYON_NUM_THREADS"] = "20000"

# Capture the warning emitted by Turbovec when it caps the value

with warnings.catch_warnings(record=True) as w:
    warnings.simplefilter("always")
    import turbovec          # the warning is generated during import

    assert any("exceeds turbovec's thread cap" in str(warn.message) for warn in w)

```

## Key Source Files for Understanding Threading

| File | Purpose |
|------|---------|
| [`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) | Python-side initialization that reads `RAYON_NUM_THREADS`, applies the cap, and builds the process-local Rayon pool. |
| [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) | Core search implementation; receives the thread count (`n_threads`) from the Rayon pool and drives the block-parallel scheduling. |
| [`README.md`](https://github.com/RyanCodrai/turbovec/blob/main/README.md) (repository root) | General usage documentation, including a brief note on environment-variable configuration for parallelism. |

## Performance Considerations

- **Setting too many threads** wastes memory and can cause context-switching overhead. The ×4 cap exists precisely to prevent pathological over-subscription.
- **Setting too few threads** underutilizes modern multicore hardware, especially on retrieval workloads of high query counts.
- For single-query or low-latency workloads, a thread count equal to your physical core count is usually optimal. For batch retrieval with many queries, pushing closer to the cap can give better throughput.

## Summary

- Turbovec controls parallel search thread count exclusively through the `RAYON_NUM_THREADS` environment variable.
- The value is read once during `import turbovec`, clamped to a cap of `4 × hardware_parallelism` (fallback 1024).
- Over-range requests produce a Python `RuntimeWarning` but still execute safely.
- The core thread pool in [`turbovec/src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/search.rs) reads the count at search time, making the whole pipeline consistent.
- For reproducibility, always set the variable explicitly in your orchestration scripts or CI environment before importing turbovec.

## Frequently Asked Questions

### How do I check the current thread count in turbovec?

While turbovec does not expose a direct public API for the thread count, the `desired_threads()` function in [`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) calculates the effective value. You can inspect the warning you receive on capping, or simply set a known value and trust the clamp.

### Does changing `RAYON_NUM_THREADS` after import affect running searches?

No. The environment variable is read once during the import-time initialization of the Rayon pool. Changing it afterward in the same Python process has no effect on future tasks. You must restart the interpreter to apply a new setting.

### What happens if I set `RAYON_NUM_THREADS` to a non-numeric value?

If the value cannot be parsed as a `usize`, `desired_threads()` falls back to the OS-detected hardware parallelism. No panic occurs; the default is a safe and sensible choice for most systems.

### The cap seems low for my 64-core server. Can I increase it?

The cap is hard-coded in [`turbovec-python/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) as four times the detected hardware parallelism. Since this value is compiled in, you cannot raise it from Python. If your workload genuinely requires more threads, you would need to fork the repository and adjust the `rayon_thread_cap` constant before rebuilding the bindings.