How to Configure Threading for Parallel Search in Turbovec Using `RAYON_NUM_THREADS`

TLDR: Set the RAYON_NUM_THREADS environment variable before importing turbovec to control the Rayon thread pool size used for parallel vector search; the value is clamped to a maximum of 4× hardware parallelism, with a fallback cap of 1024.

Turbovec is a high-performance Rust-based vector search library with Python bindings that leverages the Rayon crate for CPU-bound parallelism. According to the source code in RyanCodrai/turbovec, the number of worker threads in the process-local Rayon pool is entirely controlled by the RAYON_NUM_THREADS environment variable. This guide explains exactly how the threading configuration works under the hood and provides runnable Python examples for both manual and automatic thread management.

Where Rayon Threading Is Initialized

Turbovec configures its Rayon pool during the Python binding initialization. The critical logic lives in turbovec-python/src/lib.rs, specifically in the init_rayon_pool function. When you run import turbovec, the following sequence executes:

  1. The code reads RAYON_NUM_THREADS from the environment and parses it as a usize.
  2. It clamps the requested value against a maximum detected cap (rayon_thread_cap).
  3. If the requested count exceeds the cap, a RuntimeWarning is emitted through Python's warnings module.
  4. The actual pool size is returned by the desired_threads() function, which falls back to OS-reported parallelism when the environment variable is absent or unparseable.

The core search operation in turbovec/src/search.rs then queries the pool size at runtime, ensuring every parallel search respects the configured thread count.

The Thread Cap Logic

Turbovec does not allow unlimited threads. The source code in turbovec-python/src/lib.rs defines the cap as four times the detected hardware parallelism (lines 1906–1918). If the system's hardware parallelism cannot be determined, the cap falls back to 1024 threads.

Here's how the clamping behaves in practice:

  • If you request 8 threads on a 4-core machine, you get 8 threads (assuming the cap of 16 or more is not exceeded).
  • If you request 200 threads on a 32-core machine, the cap of 128 (4 × 32) reduces the pool to 128 threads.
  • If you request an absurd value like 20000, the pool is capped at the maximum and a warning informs you of the reduction.

How Parallel Search Uses the Thread Count

During a search operation, turbovec's core scheduling logic derives the worker thread count from the Rayon pool. In turbovec/src/search.rs (lines 1330–1336), the variable n_threads is passed to the block-parallel search engine. This means:

  • Every search invocation uses the same global thread pool.
  • The parallelism is applied across data blocks, making heavy retrieval tasks faster on multi-core systems.
  • Changes to the environment variable require a new Python process to take effect.

Setting the Thread Count: Practical Examples

Example 1: Explicit Thread Configuration

Set RAYON_NUM_THREADS before importing turbovec to guarantee a fixed pool size:

import os

# Use 8 worker threads for the Rayon pool (will be clamped to <= cap)

os.environ["RAYON_NUM_THREADS"] = "8"

import turbovec  # init_rayon_pool runs here and reads the env var

# Now run a parallel search; the pool will have exactly 8 threads

vectors = ...          # (n_vectors, dim) float32 ndarray

queries = ...          # (n_queries, dim) float32 ndarray

k = 10                 # retrieve top-k results

results = turbovec.search(vectors, queries, k)

Example 2: Automatic Hardware Detection

If you omit the environment variable, turbovec falls back to the OS-reported parallelism. This is ideal for ensuring portability across machines:

import turbovec

# Turbovec will use the hardware-detected parallelism (×4 cap) automatically

vectors = ...
queries = ...
results = turbovec.search(vectors, queries, k=5)

Example 3: Handling an Over-Large Request

The clamping produces a warning when the requested value exceeds the cap. You can capture and assert on that warning:

import os, warnings

# Request an unrealistically high thread count

os.environ["RAYON_NUM_THREADS"] = "20000"

# Capture the warning emitted by Turbovec when it caps the value

with warnings.catch_warnings(record=True) as w:
    warnings.simplefilter("always")
    import turbovec          # the warning is generated during import

    assert any("exceeds turbovec's thread cap" in str(warn.message) for warn in w)

Key Source Files for Understanding Threading

File Purpose
turbovec-python/src/lib.rs Python-side initialization that reads RAYON_NUM_THREADS, applies the cap, and builds the process-local Rayon pool.
turbovec/src/search.rs Core search implementation; receives the thread count (n_threads) from the Rayon pool and drives the block-parallel scheduling.
README.md (repository root) General usage documentation, including a brief note on environment-variable configuration for parallelism.

Performance Considerations

  • Setting too many threads wastes memory and can cause context-switching overhead. The ×4 cap exists precisely to prevent pathological over-subscription.
  • Setting too few threads underutilizes modern multicore hardware, especially on retrieval workloads of high query counts.
  • For single-query or low-latency workloads, a thread count equal to your physical core count is usually optimal. For batch retrieval with many queries, pushing closer to the cap can give better throughput.

Summary

  • Turbovec controls parallel search thread count exclusively through the RAYON_NUM_THREADS environment variable.
  • The value is read once during import turbovec, clamped to a cap of 4 × hardware_parallelism (fallback 1024).
  • Over-range requests produce a Python RuntimeWarning but still execute safely.
  • The core thread pool in turbovec/src/search.rs reads the count at search time, making the whole pipeline consistent.
  • For reproducibility, always set the variable explicitly in your orchestration scripts or CI environment before importing turbovec.

Frequently Asked Questions

How do I check the current thread count in turbovec?

While turbovec does not expose a direct public API for the thread count, the desired_threads() function in turbovec-python/src/lib.rs calculates the effective value. You can inspect the warning you receive on capping, or simply set a known value and trust the clamp.

Does changing RAYON_NUM_THREADS after import affect running searches?

No. The environment variable is read once during the import-time initialization of the Rayon pool. Changing it afterward in the same Python process has no effect on future tasks. You must restart the interpreter to apply a new setting.

What happens if I set RAYON_NUM_THREADS to a non-numeric value?

If the value cannot be parsed as a usize, desired_threads() falls back to the OS-detected hardware parallelism. No panic occurs; the default is a safe and sensible choice for most systems.

The cap seems low for my 64-core server. Can I increase it?

The cap is hard-coded in turbovec-python/src/lib.rs as four times the detected hardware parallelism. Since this value is compiled in, you cannot raise it from Python. If your workload genuinely requires more threads, you would need to fork the repository and adjust the rayon_thread_cap constant before rebuilding the bindings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →