# What Kind of Background Jobs Does Colibri Support?

> Discover Colibri's background job support including Windows Job Objects for process management and a Background I/O Loader Pool for efficient MoE shard loading during inference.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Colibri supports two primary background job types: Windows Job Objects that enforce kill-on-close semantics for subprocess lifecycle management, and a bounded Background I/O Loader Pool that asynchronously loads Mixture-of-Experts (MoE) shards while inference continues.**

Colibri is an inference engine designed to run workloads in trusted-execution environments without auxiliary tasks interfering with the main forward pass. To achieve this isolation and performance, the JustVugg/colibri repository implements distinct background job systems for process management and asynchronous data loading. These mechanisms ensure reliable resource cleanup on Windows and non-blocking I/O operations across POSIX systems including Linux and macOS.

## Windows Job Objects for Process Lifecycle Management

On Windows, Colibri leverages the **Windows Job Object** API with the `KILL_ON_JOB_CLOSE` flag to guarantee that child processes terminate automatically when the parent handle closes. This prevents orphaned worker processes from consuming GPU or CPU resources after the main inference engine exits.

### Implementation in c/openai_server.py

The function `_win_kill_on_close_job()` in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) creates a job object and configures it with the `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` (0x00002000) extended limit information. When the parent process assigns a child server process to this job object, Windows ensures the child receives a termination signal as soon as the parent's job handle is closed.

```python

# c/openai_server.py – simplified excerpt

import ctypes
from ctypes import wintypes

def _win_kill_on_close_job(pid: int):
    """Create a Windows job object that kills the process when the handle is closed."""
    k32 = ctypes.WinDLL("kernel32", use_last_error=True)

    # Create the job object

    job = k32.CreateJobObjectW(None, None)
    if not job:
        return None

    # Set the flag so that closing the handle kills the child

    info = ctypes.c_ulong(0x00002000)          # JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE

    k32.SetInformationJobObject(job,
        9,                                     # JobObjectExtendedLimitInformation

        ctypes.byref(info), ctypes.sizeof(info))

    # Assign the child process to the job

    handle = k32.OpenProcess(0x001F0FFF, False, pid)
    if not handle:
        k32.CloseHandle(job)
        return None
    k32.AssignProcessToJobObject(job, handle)
    k32.CloseHandle(handle)
    return job   # Caller must eventually `CloseHandle(job)`

```

When starting a CUDA-accelerated server or other child process, the parent calls `_win_kill_on_close_job(child.pid)` and stores the returned handle. If the parent crashes or exits normally and closes the handle, Windows immediately terminates the entire job tree, eliminating resource leaks.

## Background I/O Loader Pool for Asynchronous Model Loading

For loading model components without stalling the inference loop, Colibri implements a **bounded background thread pool** that handles disk and network I/O. This system is particularly critical for Mixture-of-Experts (MoE) models where individual "expert" shards may need loading on-demand while the main thread processes the current token batch.

### Loading MoE Experts Without Blocking Inference

The loader pool is instantiated in [`c/v4_dsml.py`](https://github.com/JustVugg/colibri/blob/main/c/v4_dsml.py) using Python's `threading` module and a bounded `Queue` to cap concurrent I/O operations. Worker threads dequeue load requests, fetch the required expert shards from storage or remote sources, and register them for use by the inference engine.

```python

# c/v4_dsml.py – illustration of the loader pool

import threading
from queue import Queue

def _loader_worker(task_q: Queue):
    while True:
        job = task_q.get()
        if job is None:               # sentinel → shutdown

            break
        load_expert(job)              # heavy I/O (disk or network)

        task_q.task_done()

# Create a bounded pool (e.g. 4 concurrent loaders)

LOADER_POOL_SIZE = 4
loader_q = Queue(maxsize=LOADER_POOL_SIZE)
for _ in range(LOADER_POOL_SIZE):
    threading.Thread(target=_loader_worker,
                     args=(loader_q,),
                     daemon=True).start()

# Somewhere in the main inference loop:

if missing_expert := find_missing_expert():
    loader_q.put(missing_expert)      # enqueue background load

```

This architecture allows the forward pass to remain responsive. While the main inference thread continues processing, the background workers fetch missing experts. The bounded queue size protects the system from I/O over-subscription that could otherwise degrade latency.

### Tuning and Configuration

The behavior and performance characteristics of this pool are documented in [`docs/tuning.md`](https://github.com/JustVugg/colibri/blob/main/docs/tuning.md), which explains how to adjust thread counts and queue depths based on storage bandwidth and model parallelism settings. Concrete usage patterns appear in [`docs/experiments/e5-metal-residency-set.md`](https://github.com/JustVugg/colibri/blob/main/docs/experiments/e5-metal-residency-set.md), demonstrating how background loader threads register new experts while other MoE blocks execute concurrently.

## Cross-Platform Background Job Architecture

Colibri selects the appropriate background job mechanism based on the host operating system:

- **Windows**: Uses **Job Objects** (`KILL_ON_JOB_CLOSE`) for reliable subprocess termination and the **Background I/O Loader Pool** for asynchronous model component loading.
- **Linux/macOS**: Relies exclusively on the **Background I/O Loader Pool** for async loading, using POSIX threads to keep the inference engine unblocked during heavy I/O operations.

Both systems execute concurrently with the main inference loop, ensuring that trusted-execution environments maintain strict resource boundaries while maximizing throughput.

## Summary

- **Windows Job Objects** provide automatic cleanup of child processes via `KILL_ON_JOB_CLOSE` semantics, implemented in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) to prevent orphaned workers.
- **Background I/O Loader Pools** handle asynchronous loading of MoE expert shards and other model assets using a bounded thread queue in [`c/v4_dsml.py`](https://github.com/JustVugg/colibri/blob/main/c/v4_dsml.py).
- The loader pool caps concurrent operations through a configurable size limit, protecting system resources from over-subscription.
- These mechanisms operate across Windows, Linux, and macOS, with Windows utilizing both job objects and loader pools while POSIX systems use loader pools exclusively.

## Frequently Asked Questions

### How does Colibri prevent orphaned subprocesses on Windows?

Colibri creates a Windows Job Object with the `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` flag in `_win_kill_on_close_job()` within [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py). When the parent process handle closes—whether through normal exit or crash—the operating system automatically terminates all processes assigned to that job object, ensuring no orphaned model servers consume GPU resources.

### What is the purpose of the background loader pool in Colibri?

The background loader pool enables **asynchronous loading of Mixture-of-Experts (MoE) shards** and other model assets without blocking the main inference thread. While the foreground processes the current token batch, background threads fetch missing experts from disk or network, maintaining low latency for the active forward pass.

### Where is the background I/O pool size configured?

The pool size is defined by the `LOADER_POOL_SIZE` constant in [`c/v4_dsml.py`](https://github.com/JustVugg/colibri/blob/main/c/v4_dsml.py), which sets both the thread count and the maximum queue depth. Performance tuning guidance for adjusting this value based on storage bandwidth appears in [`docs/tuning.md`](https://github.com/JustVugg/colibri/blob/main/docs/tuning.md), and experimental validation is documented in [`docs/experiments/e5-metal-residency-set.md`](https://github.com/JustVugg/colibri/blob/main/docs/experiments/e5-metal-residency-set.md).

### Does Colibri support background jobs on macOS?

Yes. While macOS does not use Windows Job Objects, it fully supports the **Background I/O Loader Pool** using POSIX threads. The implementation in [`c/v4_dsml.py`](https://github.com/JustVugg/colibri/blob/main/c/v4_dsml.py) uses standard Python `threading` modules compatible with macOS, allowing the same non-blocking expert loading behavior available on Linux and Windows.