# How CAPEv2's Scheduler Manages Analysis Tasks and VM Resource Allocation

> Discover how CAPEv2's scheduler manages analysis tasks and VM resource allocation, orchestrating automated malware analysis by matching tasks to VMs and enforcing resource limits.

- Repository: [Kevin O'Reilly/capev2](https://github.com/kevoreilly/capev2)
- Tags: internals
- Published: 2026-03-05

---

**CAPEv2's Scheduler continuously polls the database for pending analysis tasks, matches them to available virtual machines through the MachineryManager, and enforces resource limits via semaphores and configuration settings to orchestrate automated malware analysis.**

The CAPEv2 sandbox relies on a robust scheduling engine to coordinate malware analysis at scale. Located in [`lib/cuckoo/core/scheduler.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/scheduler.py), the **CAPEv2 scheduler** serves as the central orchestrator that bridges the database queue with virtual machine infrastructure, ensuring efficient VM resource allocation while preventing system overload.

## The Main Execution Loop

The scheduler's lifecycle begins with `Scheduler.start()`, which initializes the main processing loop. This method sets `self.loop_state = LoopState.RUNNING` and enters a continuous cycle that repeatedly invokes `do_main_loop_work()` until a shutdown signal is received.

After each iteration, the scheduler sleeps for a configurable delay returned by `do_main_loop_work()`, preventing CPU thrashing while maintaining responsiveness to new tasks. This sleep mechanism, found at lines 71-74 of [`lib/cuckoo/core/scheduler.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/scheduler.py), allows the system to dynamically adjust polling frequency based on current load.

## Task Selection and Prioritization

Inside `do_main_loop_work()`, the scheduler first validates global resource constraints before attempting to assign work. These checks include verifying maximum analysis counts, available disk space, and cleaning tasks that have exceeded timeout thresholds.

The scheduler then queries the **MachineryManager** to determine if the system has reached its concurrent VM limit via `self.machinery_manager.running_machines_max_reached()`. If capacity remains available, the scheduler proceeds to `find_next_serviceable_task()`, which delegates to one of two specialized methods:

- `find_pending_task_to_service()` – Identifies tasks requiring virtual machine execution
- `find_pending_task_not_requiring_machinery()` – Identifies tasks that can run without VM resources

## Matching Tasks to Virtual Machines

When processing VM-dependent tasks, `find_pending_task_to_service()` iterates over pending analysis entries and queries `MachineryManager.find_machine_to_service_task(task)` to locate compatible infrastructure. This matching process considers task requirements such as platform, architecture, and available machine tags.

If no suitable VM exists for a given task, the scheduler consults the `cuckoo.fail_unserviceable` configuration option. When enabled, the task is immediately marked as failed; otherwise, it remains in the queue as unserviceable until matching infrastructure becomes available. This logic appears in lines 236-255 of [`lib/cuckoo/core/scheduler.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/scheduler.py).

## VM Lifecycle Management

Upon identifying a compatible machine, the scheduler instantiates a new `AnalysisManager` object, passing the `machinery_manager` reference to handle VM operations. The `AnalysisManager.start()` method subsequently calls `MachineryManager.start_machine(machine)` to initiate the virtual machine.

The `start_machine()` method in [`lib/cuckoo/core/machinery_manager.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/machinery_manager.py) implements sophisticated resource locking:

1. For auto-scaling backends using **ScalingBoundedSemaphore**, it releases a permit if the semaphore value is zero
2. It acquires the appropriate lock mechanism (Lock, BoundedSemaphore, or ScalingBoundedSemaphore)
3. Finally, it invokes the machinery's `start()` method to launch the VM

## Resource Allocation Controls

The CAPEv2 scheduler enforces multiple layers of resource protection to prevent infrastructure exhaustion:

**Maximum Concurrent VMs** – The `MachineryManager.running_machines_max_reached()` method compares active machine counts against configured limits, preventing the scheduler from over-committing resources.

**Auto-scaling Semaphore** – When `cuckoo.scaling_semaphore` is enabled, `MachineryManager.create_machine_lock()` instantiates a `ScalingBoundedSemaphore` with an upper limit derived from `machines_limit`. A background maintenance thread (`thr_maintain_scaling_bounded_semaphore`) periodically adjusts the semaphore limit and monitors for starvation conditions.

**Stuck VM Detection** – Each main loop iteration inspects running analyses for tasks exceeding their allocated runtime. When `duration > max_runtime`, the scheduler terminates the VM and logs diagnostic stack traces, preventing indefinite resource blocking.

## Graceful Shutdown

When the scheduler receives a stop signal, it transitions to `LoopState.STOPPING` and invokes `wait_for_running_analyses_to_finish()`. This method blocks until all active `AnalysisManager` threads complete their current tasks, ensuring no analyses are aborted mid-execution before the scheduler becomes fully inactive.

## Code Examples

The following snippets demonstrate how developers can interact with the CAPEv2 scheduler components programmatically:

```python
from lib.cuckoo.core.scheduler import Scheduler
from queue import Queue

# Create a scheduler that will process up to 10 analyses before exiting

sched = Scheduler(maxcount=10)

# Run the scheduler in the current thread (blocking)

error_q = Queue()
sched.start()          # starts the main loop, which internally calls do_main_loop_work()

```

To manually request VM resources for a specific task from a custom script:

```python
from lib.cuckoo.core.machinery_manager import MachineryManager
from lib.cuckoo.core.data.task import Task

mm = MachineryManager()
task = Task.get(task_id)        # obtain a Task object from the DB

machine = mm.find_machine_to_service_task(task)
if machine:
    mm.start_machine(machine)  # launches the VM according to the configured lock/semaphore

```

## Summary

- The **CAPEv2 scheduler** operates through a continuous polling loop in `Scheduler.start()`, delegating work to `do_main_loop_work()` while respecting configurable sleep intervals.
- **Task selection** involves checking global limits, querying the **MachineryManager** for capacity, and matching pending tasks to compatible VMs via `find_machine_to_service_task()`.
- **VM resource allocation** uses sophisticated locking mechanisms including `ScalingBoundedSemaphore` for auto-scaling environments, with background threads monitoring resource starvation.
- **Safety mechanisms** include stuck VM detection, graceful shutdown handling, and the `cuckoo.fail_unserviceable` configuration for handling tasks that cannot be matched to infrastructure.

## Frequently Asked Questions

### How does CAPEv2 prevent the scheduler from over-committing VM resources?

The scheduler consults `MachineryManager.running_machines_max_reached()` before assigning new tasks, which compares active machine counts against the configured `max_machines_count`. Additionally, the system uses semaphores (including `ScalingBoundedSemaphore` for auto-scaling backends) to enforce hard limits on concurrent VM operations.

### What happens when CAPEv2 cannot find a suitable VM for a pending analysis task?

When `find_machine_to_service_task()` returns no compatible machine, the scheduler checks the `cuckoo.fail_unserviceable` configuration option. If enabled, the task is immediately marked as failed and removed from the queue. If disabled, the task remains pending as "unserviceable" until matching infrastructure becomes available or an operator intervenes.

### How does CAPEv2 handle stuck or hung virtual machines during analysis?

Each iteration of the main scheduler loop inspects running analyses for timeout violations. When a task's duration exceeds `max_runtime`, the scheduler terminates the associated VM and logs a diagnostic stack trace. This prevents indefinitely blocked resources and ensures the system can recover from misbehaving guest environments.

### Can the CAPEv2 scheduler operate with auto-scaling VM backends?

Yes, when `cuckoo.scaling_semaphore` is enabled, the `MachineryManager` instantiates a `ScalingBoundedSemaphore` instead of a standard lock. A dedicated background thread (`thr_maintain_scaling_bounded_semaphore`) periodically adjusts the semaphore's capacity based on the current `machines_limit`, allowing the scheduler to dynamically scale VM resources up or down according to backend availability.