# How to Develop and Implement Custom Processing Modules for CAPEv2

> Learn to develop and implement custom processing modules for CAPEv2. Create Python classes inheriting from the Processing base class to transform artifacts into structured data. Explore the kevoreilly/capev2 repository for exam...

- Repository: [Kevin O'Reilly/capev2](https://github.com/kevoreilly/capev2)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Custom CAPEv2 processing modules are Python classes inheriting from the `Processing` base class that transform raw analysis artifacts into structured data by implementing a `run()` method and placing the file in `modules/processing/`.**

CAPEv2 uses a modular pipeline to analyze malware samples, and understanding how to develop and implement custom processing modules for CAPEv2 allows you to extend its capabilities beyond the default behavior analysis. The framework automatically discovers and executes these modules, merging their output into the final analysis report without requiring modifications to the core scheduler.

## Understanding the CAPEv2 Processing Pipeline Architecture

### The Processing Base Class

Every processing module must inherit from the abstract `Processing` class defined in [`lib/cuckoo/common/abstracts.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/common/abstracts.py). This base class provides the essential API contract that the orchestration engine expects, including helper methods for path resolution, task metadata access, and statistics tracking.

The `Processing` class defines three critical setter methods that the framework calls before `run()`:
- `set_path()` – Injects the absolute path to the analysis directory
- `set_task()` – Provides the task dictionary containing target file, platform, and options
- `set_options()` – Supplies module-specific configuration from [`processing.conf`](https://github.com/kevoreilly/capev2/blob/main/processing.conf)

### How RunProcessing Orchestrates Modules

The `RunProcessing` class in [`lib/cuckoo/core/plugins.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/plugins.py) manages the execution lifecycle. When a task completes, the scheduler instantiates `RunProcessing` and invokes its `process()` method. This method iterates over every Python file in `modules/processing/`, imports classes inheriting from `Processing`, checks the configuration for an `enabled` flag, and executes the `run()` method.

Timing data is automatically captured and stored in `results["statistics"]["processing"]` without requiring explicit instrumentation in your module.

## Step-by-Step Guide to Creating Custom Processing Modules

### Step 1: Create the Module File in modules/processing/

Navigate to the `modules/processing/` directory and create a new Python file. The filename should be descriptive but does not need to match the class name. For example, [`pe_counter.py`](https://github.com/kevoreilly/capev2/blob/main/pe_counter.py) or [`network_enhancer.py`](https://github.com/kevoreilly/capev2/blob/main/network_enhancer.py).

### Step 2: Inherit from the Processing Base Class

Import the base class and define your module:

```python
from lib.cuckoo.common.abstracts import Processing

class PeCounter(Processing):
    """Counts PE files extracted during analysis."""
    pass

```

### Step 3: Implement the run() Method

The `run()` method performs your analysis logic and must return a dictionary. Set `self.key` to define where the data appears in the final report:

```python
def run(self):
    self.key = "pe_counter"
    # Access analysis path via self.analysis_path

    # Access task info via self.task

    # Access config options via self.options

    return {"count": 0, "files": []}

```

### Step 4: Configure the Module in processing.conf

Add a section to [`conf/processing.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/processing.conf) (or the equivalent JSON/YAML configuration):

```ini
[pe_counter]
enabled = true
on_demand = false

```

The section name must match your class name in lowercase. Set `enabled = true` to activate the module.

## Practical Code Examples for CAPEv2 Processing Modules

### Example 1: Counting PE Files in Analysis Artifacts

This complete module demonstrates accessing the analysis directory and returning structured data:

```python

# modules/processing/pe_counter.py

import os
from lib.cuckoo.common.abstracts import Processing

class PeCounter(Processing):
    """Counts PE files in the `files` directory of an analysis."""
    def run(self):
        # Define the key for results aggregation

        self.key = "pe_counter"
        
        # Construct path to extracted files

        files_dir = os.path.join(self.analysis_path, "files")
        
        # Safely handle missing directories

        if not os.path.exists(files_dir):
            return {"total_pe": 0, "pe_list": []}
        
        pe_files = [f for f in os.listdir(files_dir) if f.lower().endswith(".exe")]
        
        return {"total_pe": len(pe_files), "pe_list": pe_files}

```

The `self.analysis_path` attribute is automatically populated by the framework via `set_path()`.

### Example 2: Accessing Task Metadata and Previous Results

This module demonstrates reading from `self.task` and `self.results` to build composite reports:

```python

# modules/processing/report_summary.py

from lib.cuckoo.common.abstracts import Processing

class ReportSummary(Processing):
    """Creates a summary using data from earlier processing modules."""
    def run(self):
        self.key = "summary"
        
        # Access task metadata

        target = self.task.get("target", "unknown")
        platform = self.task.get("platform", "windows")
        
        # Access results from other modules

        behavior = self.results.get("behavior", {})
        processes = behavior.get("processes", [])
        num_procs = len(processes)
        
        # Access memory module results if available

        memory_data = self.results.get("memory", {})
        mem_usage = memory_data.get("mem_usage", [])
        avg_mem = sum(mem_usage) / len(mem_usage) if mem_usage else 0
        
        return {
            "target": target,
            "platform": platform,
            "process_count": num_procs,
            "average_memory_mb": round(avg_mem, 2),
            "generated_by": __class__.__name__
        }

```

### Example 3: Debugging Modules with RunProcessing

Test your module without running a full analysis using the `RunProcessing` class directly:

```python
>>> from lib.cuckoo.core.plugins import RunProcessing
>>> 
>>> # Create mock task and results structure

>>> task = {"id": 42, "platform": "windows", "target": "sample.exe"}
>>> results = {"statistics": {"processing": []}}
>>> 
>>> # Initialize the runner

>>> runner = RunProcessing(task, results)
>>> 
>>> # Import and test your custom module

>>> from modules.processing.pe_counter import PeCounter
>>> runner.process(PeCounter)
>>> 
>>> # Inspect the results

>>> results['pe_counter']
{'total_pe': 3, 'pe_list': ['a.exe', 'b.exe', 'c.exe']}

```

This approach allows rapid iteration without submitting samples to the sandbox.

## Advanced Configuration and Testing Strategies

### Enabling On-Demand Processing

CAPEv2 supports conditional execution via the `on_demand` configuration flag. When set to `true`, the module only executes when explicitly requested rather than during every analysis:

```ini
[pe_counter]
enabled = true
on_demand = true

```

In your module, check this flag via `self.options.get("on_demand")` to implement conditional logic. This pattern is used by the `url_analysis` module to avoid unnecessary API calls to VirusTotal unless specifically enabled.

### Unit Testing with ProcessingMock

The repository includes testing utilities in [`tests/processor_tests.py`](https://github.com/kevoreilly/capev2/blob/main/tests/processor_tests.py). Use the `ProcessingMock` pattern to instantiate your module with controlled inputs:

```python

# tests/test_pe_counter.py

import unittest
from modules.processing.pe_counter import PeCounter
from tests.processor_tests import ProcessingMock

class TestPeCounter(unittest.TestCase):
    def test_empty_directory(self):
        # Setup mock with temporary path

        mock = ProcessingMock()
        module = PeCounter()
        module.set_path("/tmp/empty_analysis")
        module.set_task({"id": 1, "target": "test.exe"})
        module.set_options({"enabled": "true"})
        
        result = module.run()
        self.assertEqual(result["total_pe"], 0)

```

This approach isolates your logic from the CAPEv2 runtime while verifying correct behavior.

## Summary

- **Inherit from `Processing`**: All custom modules must subclass the abstract base class in [`lib/cuckoo/common/abstracts.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/common/abstracts.py) and implement the `run()` method.
- **Place files in `modules/processing/`**: The `RunProcessing` engine automatically discovers Python files in this directory; no manual registration is required.
- **Return dictionaries**: The `run()` method must return a dict that gets merged into the global results under the key specified by `self.key`.
- **Configure via [`processing.conf`](https://github.com/kevoreilly/capev2/blob/main/processing.conf)**: Enable modules by adding a lowercase section matching the class name with `enabled = true`.
- **Leverage built-in helpers**: Use `self.analysis_path`, `self.task`, and `self.options` to access analysis data, and rely on automatic timing statistics collection.

## Frequently Asked Questions

### What is the minimum required method to implement in a CAPEv2 processing module?

You must implement the `run()` method. This method receives no arguments (all context is available via `self` attributes set by the framework) and must return a dictionary containing your analysis results. The base class `Processing` in [`lib/cuckoo/common/abstracts.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/common/abstracts.py) defines this as an abstract method, so the framework will raise an error if you omit it.

### Where does CAPEv2 look for custom processing modules?

CAPEv2 scans the `modules/processing/` directory for Python files containing classes that inherit from `Processing`. The `RunProcessing` class in [`lib/cuckoo/core/plugins.py`](https://github.com/kevoreilly/capev2/blob/main/lib/cuckoo/core/plugins.py) handles this discovery automatically during initialization. You do not need to import or register your module manually; simply placing the file in the correct directory with the proper inheritance pattern makes it available to the framework.

### How do I access analysis results from other processing modules in my custom module?

Access the global results dictionary via `self.results`, which the framework populates with data from previously executed modules. For example, to read behavioral data from the `behavior` module, use `self.results.get("behavior", {})`. To access the current task's metadata (such as target file or platform), use `self.task`. These attributes are injected by `RunProcessing` before your `run()` method executes.

### Can I disable a processing module without removing the code?

Yes. Add a configuration section to [`conf/processing.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/processing.conf) (or the equivalent JSON/YAML configuration) with the same name as your class in lowercase, and set `enabled = false`. The `RunProcessing` engine checks this flag before instantiating your module. You can also implement an `on_demand` option to allow conditional execution only when explicitly requested, which is useful for expensive operations like third-party API lookups.