# How to Configure CAPEv2 for Integration with Cloud-Based Virtual Machine Providers

> Integrate CAPEv2 with AWS Azure or GCP by installing SDKs and configuring cuckoo.conf Access cloud-based virtual machines for advanced analysis.

- Repository: [Kevin O'Reilly/capev2](https://github.com/kevoreilly/capev2)
- Tags: how-to-guide
- Published: 2026-03-05

---

**To configure CAPEv2 for integration with cloud-based virtual machine providers like AWS, Azure, or GCP, install the provider-specific Python SDK, set `machinery = aws`, `machinery = az`, or `machinery = gcp` in [`conf/cuckoo.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/cuckoo.conf), and copy the default configuration template from `conf/default/` to configure credentials, networking, and autoscaling parameters.**

The CAPEv2 malware sandbox (maintained in the `kevoreilly/capev2` repository) supports elastic analysis infrastructure by provisioning virtual machines on Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) instead of local hypervisors. This guide explains how to configure the machinery modules located in `modules/machinery/` to orchestrate cloud-based analysis VMs, referencing the actual implementation in [`aws.py`](https://github.com/kevoreilly/capev2/blob/main/aws.py), [`az.py`](https://github.com/kevoreilly/capev2/blob/main/az.py), and the configuration schemas provided in the source tree.

## Prerequisites for Cloud VM Integration

Before enabling cloud machinery, verify that your CAPE host meets the following requirements for API access and network connectivity.

### Install Provider-Specific Python Packages

Each cloud provider requires a distinct set of Python libraries. Install them using pip within your CAPE virtual environment:

- **AWS**: Requires `boto3` for EC2 API interaction.
- **Azure**: Requires `azure-identity`, `azure-mgmt-compute`, `azure-mgmt-network`, and `msrestazure` for VM Scale Set management.
- **GCP**: Requires `google-auth` and `google-api-python-client` for Compute Engine operations.

```bash

# AWS dependencies

poetry run pip install boto3

# Azure dependencies  

poetry run pip install azure-identity msrest msrestazure azure-mgmt-compute azure-mgmt-network

# GCP dependencies

poetry run pip install google-auth google-api-python-client

```

### Configure Cloud Credentials

The machinery modules support multiple authentication methods. Avoid hardcoding secrets in production by using IAM roles or managed identities where possible.

- **AWS**: Provide `aws_access_key_id` and `aws_secret_access_key` in [`conf/aws.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/aws.conf), or leave these fields blank to use an **IAM role** attached to the host instance (IMDS).
- **Azure**: Configure a **Service Principal** by setting `client_id`, `secret`, and `tenant` in [`conf/az.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/az.conf). Alternatively, specify `certificate_path` and `certificate_password` for certificate-based authentication.
- **GCP**: Set `service_account_path` to a JSON key file in [`conf/gcp.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/gcp.conf), or set `running_in_gcp = true` to use the default Compute Engine service account when CAPE runs inside GCP.

### Verify Network Accessibility

The CAPE result server must accept connections from cloud VMs on ports **2042**, **8000**, and **8090**. Configure your cloud security groups (AWS), network security groups (Azure), or firewall rules (GCP) to allow inbound traffic from the analysis subnet to these ports on the CAPE host.

## Enable Cloud Machinery in CAPE

Global machinery selection is controlled in the main CAPE configuration file.

### Select the Machinery Module

Edit [`conf/cuckoo.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/cuckoo.conf) (or `conf/default/cuckoo.conf.default` if merging configurations) to specify the cloud provider:

```ini
[cuckoo]

# Valid options: aws, az, gcp

machinery = aws

```

### Copy Default Provider Configurations

Copy the template configuration for your selected provider from the defaults directory to the active configuration directory:

```bash
cp conf/default/aws.conf.default conf/aws.conf
cp conf/default/az.conf.default conf/az.conf
cp conf/default/gcp.conf.default conf/gcp.conf

```

Edit the copied file to set region, credential, and networking parameters specific to your environment.

## Provider-Specific Configuration

Each cloud provider uses a dedicated machinery class in `modules/machinery/` that parses its respective configuration file for VM lifecycle management.

### AWS EC2 Configuration

The **AWS machinery** ([`modules/machinery/aws.py`](https://github.com/kevoreilly/capev2/blob/main/modules/machinery/aws.py)) manages EC2 instances using the `boto3` SDK. It supports both static pre-provisioned instances and dynamic autoscaling pools.

Key configuration parameters in [`conf/aws.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/aws.conf):

- `region_name` and `availability_zone`: Define the AWS region (e.g., `us-east-1`) and specific AZ for volume placement.
- `machines`: Optional comma-separated list of existing instance IDs (e.g., `i-0123456789abcdef`) to use as static analysis VMs.
- `autoscale` block: Enable dynamic provisioning by setting `autoscale = yes`, `dynamic_machines_limit` (max VMs), `image_id` (AMI), `instance_type` (e.g., `t2.medium`), `subnet_id`, and `security_groups`.
- `running_machines_gap`: Desired number of idle VMs ready for immediate task assignment (default is typically `1`).

When autoscaling is enabled, the `AWS` class monitors the pool size. If available VMs fall below `running_machines_gap`, it creates new instances tagged with `AUTOSCALE_CUCKOO=True` and registers them in the CAPE database. Upon task completion, autoscaled instances are terminated, while static instances are stopped.

Example [`conf/aws.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/aws.conf) snippet:

```ini
[aws]
region_name = us-east-1
availability_zone = us-east-1a

# Leave blank to use IAM role

# aws_access_key_id = YOUR_KEY_ID

# aws_secret_access_key = YOUR_SECRET

[autoscale]
autoscale = yes
dynamic_machines_limit = 3
image_id = ami-0abc1234def56789
instance_type = t2.medium
subnet_id = subnet-0a1b2c3d4e5f6g7h
security_groups = sg-0123abcd4567efgh
platform = windows
arch = x64
running_machines_gap = 1

```

### Microsoft Azure Configuration

The **Azure machinery** ([`modules/machinery/az.py`](https://github.com/kevoreilly/capev2/blob/main/modules/machinery/az.py)) orchestrates **Virtual Machine Scale Sets (VMSS)** using the Azure Management SDK. It runs a background monitoring thread (`_thr_machine_pool_monitor`) that checks pool health every `monitor_rate` seconds (default **300**).

Key configuration parameters in [`conf/az.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/az.conf):

- `region_name`: Azure region (e.g., `eastus`).
- `subscription_id`, `client_id`, `secret`, `tenant`: Service Principal credentials for the Azure SDK (`ComputeManagementClient` and `NetworkManagementClient`).
- `vnet_resource_group` and `sandbox_resource_group`: Resource groups containing the virtual network and the sandbox VMs respectively.
- `vnet` and `subnet`: Names of the existing virtual network and subnet for VM deployment.
- `scale_sets`: Comma-separated list of VM Scale Set names to manage (e.g., `cuckoo-scale-set`).
- `gallery_image_name`: Reference to an Azure Shared Image Gallery image for VM creation.
- `initial_pool_size`: Baseline number of VMs to maintain in the scale set.

The module ensures each scale set has the `AUTO_SCALE_CAPE=True` tag and adjusts capacity to match `initial_pool_size` plus any over-provision settings.

Example [`conf/az.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/az.conf) snippet:

```ini
[az]
region_name = eastus
subscription_id = 11111111-2222-3333-4444-555555555555
client_id = aaaa-bbbb-cccc-dddd-eeeeeeeeeeee
secret = <your-secret>
tenant = 66666666-7777-8888-9999-aaaaaaaaaaaa

vnet_resource_group = cape-vnet-rg
sandbox_resource_group = cape-sandbox-rg
vnet = cape-vnet
subnet = cape-subnet

scale_sets = cuckoo1

[cuckoo1]
gallery_image_name = my-cape-image
platform = windows
instance_type = Standard_D2s_v3
pool_tag = windows-pool
initial_pool_size = 2
tags = tag1,tag2

```

### Google Cloud Platform Configuration

CAPEv2 supports GCP through either a dedicated machinery module (in newer branches) or an EC2-compatible wrapper. Configuration resides in [`conf/gcp.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/gcp.conf) using parameters similar to AWS but with GCP-specific terminology.

Key configuration parameters in [`conf/gcp.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/gcp.conf):

- `zone`: Deployment zone (e.g., `europe-north2-a`).
- `project`: GCP project ID.
- `service_account_path`: Path to the JSON service account key file (omit if using the `running_in_gcp` flag).
- `running_in_gcp`: Set to `true` when the CAPE host runs on a GCE instance to leverage the attached service account.
- `machines`: Optional list of pre-existing instance names to use as static targets.
- `autoscale` block: Configure `machine_type` (e.g., `n1-standard-2`), `network`, `subnet`, `image_family`, and `dynamic_machines_limit`.

Example [`conf/gcp.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/gcp.conf) snippet:

```ini
[gcp]
zone = europe-north2-a
project = cape-gcp-project
running_in_gcp = true

# service_account_path = /opt/cape/gcp-sa.json

[autoscale]
autoscale = yes
dynamic_machines_limit = 2
machine_type = n1-standard-2
network = default
subnet = default
image_family = windows-2019
platform = windows
arch = x64
running_machines_gap = 1

```

## End-to-End AWS Deployment Example

The following commands illustrate a complete setup for AWS integration:

```bash

# 1. Install AWS SDK

poetry run pip install boto3

# 2. Enable AWS machinery in global config

sed -i 's/^machinery = .*/machinery = aws/' conf/cuckoo.conf

# 3. Copy and customize AWS configuration

cp conf/default/aws.conf.default conf/aws.conf

# Edit conf/aws.conf to set region, subnet, AMI ID, and IAM credentials or role

# 4. Restart CAPE services to load new machinery

systemctl restart cape

```

Apply the same pattern for Azure or GCP by substituting the machinery name, configuration file, and Python dependencies.

## Troubleshooting Cloud Integration Issues

| Symptom | Root Cause | Resolution |
|---------|------------|------------|
| `ImportError: No module named boto3` | AWS SDK not installed | Run `poetry run pip install boto3` |
| "Unable to import Azure packages" | Missing Azure libraries | Install dependencies listed at line 34 of [`modules/machinery/az.py`](https://github.com/kevoreilly/capev2/blob/main/modules/machinery/az.py) |
| "No credentials" errors | Missing IAM role or static keys | Attach IAM role with `AmazonEC2FullAccess` or provide keys in [`aws.conf`](https://github.com/kevoreilly/capev2/blob/main/aws.conf) |
| Result server connection timeouts | Security group rules blocking ports | Allow inbound TCP **2042**, **8000**, **8090** from the cloud subnet |
| Autoscaled VMs persist after analysis | Missing autoscale tag or disabled flag | Verify `autoscale = yes` and that the `AUTOSCALE_CUCKOO=True` tag is applied (see `_is_autoscaled` in [`aws.py`](https://github.com/kevoreilly/capev2/blob/main/aws.py)) |
| Azure pool size not adjusting | Monitoring thread failure or misconfiguration | Check `monitor_rate` (default 300s) and verify Service Principal permissions on the VM Scale Set |

## Summary

Configuring CAPEv2 for cloud-based virtual machine integration requires:

- **Installing provider SDKs**: `boto3` for AWS, `azure-mgmt-*` packages for Azure, or `google-api-python-client` for GCP.
- **Selecting machinery**: Set `machinery = aws|az|gcp` in [`conf/cuckoo.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/cuckoo.conf) and copy the default configuration template to the `conf/` directory.
- **Configuring autoscaling**: Define `dynamic_machines_limit`, `running_machines_gap`, and base image IDs to enable elastic VM pools.
- **Securing credentials**: Use IAM roles (AWS), Service Principals (Azure), or service accounts (GCP) rather than long-lived static credentials.
- **Opening firewall ports**: Ensure the CAPE result server ports **2042**, **8000**, and **8090** are accessible from cloud subnets.

With these settings, CAPEv2 automatically provisions, monitors, and terminates analysis VMs on your chosen cloud provider.

## Frequently Asked Questions

### How do I secure cloud credentials when configuring CAPEv2?

Use identity-based authentication instead of static keys. For AWS, attach an **IAM role** to the CAPE host and leave `aws_access_key_id` blank in [`conf/aws.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/aws.conf). For Azure, use a **Service Principal** with limited scope on the sandbox resource group. For GCP, set `running_in_gcp = true` to leverage the instance's attached service account, avoiding JSON key files entirely.

### What is the difference between static machines and autoscaling in CAPEv2?

**Static machines** are pre-provisioned instance IDs listed in the `machines` configuration key; CAPE starts and stops these existing VMs. **Autoscaling** dynamically creates and terminates instances based on the `running_machines_gap` setting—when the pool of idle VMs drops below this threshold, the machinery module (e.g., `AWS` class in [`modules/machinery/aws.py`](https://github.com/kevoreilly/capev2/blob/main/modules/machinery/aws.py)) launches new instances tagged with `AUTOSCALE_CUCKOO=True`.

### Why are my Azure VM Scale Sets not scaling automatically?

Verify that the Service Principal has **Contributor** access to both the `sandbox_resource_group` and the specific VM Scale Sets. Additionally, check that the `monitor_rate` (default **300** seconds) has elapsed, as the `_thr_machine_pool_monitor` thread in [`modules/machinery/az.py`](https://github.com/kevoreilly/capev2/blob/main/modules/machinery/az.py) only evaluates pool size at this interval. Incorrect `gallery_image_name` or `subnet` configurations will also prevent successful scaling.

### Can I run the CAPE management server on-premise while using cloud-based analysis VMs?

Yes. The CAPE host can reside on-premise or in a different cloud region, provided the **result server** (ports 2042, 8000, 8090) is reachable from the cloud subnet where analysis VMs execute. Configure your VPN, Direct Connect, or public IP with appropriate security group rules to allow this communication path.