How to Develop and Implement Custom Processing Modules for CAPEv2

Custom CAPEv2 processing modules are Python classes inheriting from the Processing base class that transform raw analysis artifacts into structured data by implementing a run() method and placing the file in modules/processing/.

CAPEv2 uses a modular pipeline to analyze malware samples, and understanding how to develop and implement custom processing modules for CAPEv2 allows you to extend its capabilities beyond the default behavior analysis. The framework automatically discovers and executes these modules, merging their output into the final analysis report without requiring modifications to the core scheduler.

Understanding the CAPEv2 Processing Pipeline Architecture

The Processing Base Class

Every processing module must inherit from the abstract Processing class defined in lib/cuckoo/common/abstracts.py. This base class provides the essential API contract that the orchestration engine expects, including helper methods for path resolution, task metadata access, and statistics tracking.

The Processing class defines three critical setter methods that the framework calls before run():

  • set_path() – Injects the absolute path to the analysis directory
  • set_task() – Provides the task dictionary containing target file, platform, and options
  • set_options() – Supplies module-specific configuration from processing.conf

How RunProcessing Orchestrates Modules

The RunProcessing class in lib/cuckoo/core/plugins.py manages the execution lifecycle. When a task completes, the scheduler instantiates RunProcessing and invokes its process() method. This method iterates over every Python file in modules/processing/, imports classes inheriting from Processing, checks the configuration for an enabled flag, and executes the run() method.

Timing data is automatically captured and stored in results["statistics"]["processing"] without requiring explicit instrumentation in your module.

Step-by-Step Guide to Creating Custom Processing Modules

Step 1: Create the Module File in modules/processing/

Navigate to the modules/processing/ directory and create a new Python file. The filename should be descriptive but does not need to match the class name. For example, pe_counter.py or network_enhancer.py.

Step 2: Inherit from the Processing Base Class

Import the base class and define your module:

from lib.cuckoo.common.abstracts import Processing

class PeCounter(Processing):
    """Counts PE files extracted during analysis."""
    pass

Step 3: Implement the run() Method

The run() method performs your analysis logic and must return a dictionary. Set self.key to define where the data appears in the final report:

def run(self):
    self.key = "pe_counter"
    # Access analysis path via self.analysis_path

    # Access task info via self.task

    # Access config options via self.options

    return {"count": 0, "files": []}

Step 4: Configure the Module in processing.conf

Add a section to conf/processing.conf (or the equivalent JSON/YAML configuration):

[pe_counter]
enabled = true
on_demand = false

The section name must match your class name in lowercase. Set enabled = true to activate the module.

Practical Code Examples for CAPEv2 Processing Modules

Example 1: Counting PE Files in Analysis Artifacts

This complete module demonstrates accessing the analysis directory and returning structured data:


# modules/processing/pe_counter.py

import os
from lib.cuckoo.common.abstracts import Processing

class PeCounter(Processing):
    """Counts PE files in the `files` directory of an analysis."""
    def run(self):
        # Define the key for results aggregation

        self.key = "pe_counter"
        
        # Construct path to extracted files

        files_dir = os.path.join(self.analysis_path, "files")
        
        # Safely handle missing directories

        if not os.path.exists(files_dir):
            return {"total_pe": 0, "pe_list": []}
        
        pe_files = [f for f in os.listdir(files_dir) if f.lower().endswith(".exe")]
        
        return {"total_pe": len(pe_files), "pe_list": pe_files}

The self.analysis_path attribute is automatically populated by the framework via set_path().

Example 2: Accessing Task Metadata and Previous Results

This module demonstrates reading from self.task and self.results to build composite reports:


# modules/processing/report_summary.py

from lib.cuckoo.common.abstracts import Processing

class ReportSummary(Processing):
    """Creates a summary using data from earlier processing modules."""
    def run(self):
        self.key = "summary"
        
        # Access task metadata

        target = self.task.get("target", "unknown")
        platform = self.task.get("platform", "windows")
        
        # Access results from other modules

        behavior = self.results.get("behavior", {})
        processes = behavior.get("processes", [])
        num_procs = len(processes)
        
        # Access memory module results if available

        memory_data = self.results.get("memory", {})
        mem_usage = memory_data.get("mem_usage", [])
        avg_mem = sum(mem_usage) / len(mem_usage) if mem_usage else 0
        
        return {
            "target": target,
            "platform": platform,
            "process_count": num_procs,
            "average_memory_mb": round(avg_mem, 2),
            "generated_by": __class__.__name__
        }

Example 3: Debugging Modules with RunProcessing

Test your module without running a full analysis using the RunProcessing class directly:

>>> from lib.cuckoo.core.plugins import RunProcessing
>>> 
>>> # Create mock task and results structure

>>> task = {"id": 42, "platform": "windows", "target": "sample.exe"}
>>> results = {"statistics": {"processing": []}}
>>> 
>>> # Initialize the runner

>>> runner = RunProcessing(task, results)
>>> 
>>> # Import and test your custom module

>>> from modules.processing.pe_counter import PeCounter
>>> runner.process(PeCounter)
>>> 
>>> # Inspect the results

>>> results['pe_counter']
{'total_pe': 3, 'pe_list': ['a.exe', 'b.exe', 'c.exe']}

This approach allows rapid iteration without submitting samples to the sandbox.

Advanced Configuration and Testing Strategies

Enabling On-Demand Processing

CAPEv2 supports conditional execution via the on_demand configuration flag. When set to true, the module only executes when explicitly requested rather than during every analysis:

[pe_counter]
enabled = true
on_demand = true

In your module, check this flag via self.options.get("on_demand") to implement conditional logic. This pattern is used by the url_analysis module to avoid unnecessary API calls to VirusTotal unless specifically enabled.

Unit Testing with ProcessingMock

The repository includes testing utilities in tests/processor_tests.py. Use the ProcessingMock pattern to instantiate your module with controlled inputs:


# tests/test_pe_counter.py

import unittest
from modules.processing.pe_counter import PeCounter
from tests.processor_tests import ProcessingMock

class TestPeCounter(unittest.TestCase):
    def test_empty_directory(self):
        # Setup mock with temporary path

        mock = ProcessingMock()
        module = PeCounter()
        module.set_path("/tmp/empty_analysis")
        module.set_task({"id": 1, "target": "test.exe"})
        module.set_options({"enabled": "true"})
        
        result = module.run()
        self.assertEqual(result["total_pe"], 0)

This approach isolates your logic from the CAPEv2 runtime while verifying correct behavior.

Summary

  • Inherit from Processing: All custom modules must subclass the abstract base class in lib/cuckoo/common/abstracts.py and implement the run() method.
  • Place files in modules/processing/: The RunProcessing engine automatically discovers Python files in this directory; no manual registration is required.
  • Return dictionaries: The run() method must return a dict that gets merged into the global results under the key specified by self.key.
  • Configure via processing.conf: Enable modules by adding a lowercase section matching the class name with enabled = true.
  • Leverage built-in helpers: Use self.analysis_path, self.task, and self.options to access analysis data, and rely on automatic timing statistics collection.

Frequently Asked Questions

What is the minimum required method to implement in a CAPEv2 processing module?

You must implement the run() method. This method receives no arguments (all context is available via self attributes set by the framework) and must return a dictionary containing your analysis results. The base class Processing in lib/cuckoo/common/abstracts.py defines this as an abstract method, so the framework will raise an error if you omit it.

Where does CAPEv2 look for custom processing modules?

CAPEv2 scans the modules/processing/ directory for Python files containing classes that inherit from Processing. The RunProcessing class in lib/cuckoo/core/plugins.py handles this discovery automatically during initialization. You do not need to import or register your module manually; simply placing the file in the correct directory with the proper inheritance pattern makes it available to the framework.

How do I access analysis results from other processing modules in my custom module?

Access the global results dictionary via self.results, which the framework populates with data from previously executed modules. For example, to read behavioral data from the behavior module, use self.results.get("behavior", {}). To access the current task's metadata (such as target file or platform), use self.task. These attributes are injected by RunProcessing before your run() method executes.

Can I disable a processing module without removing the code?

Yes. Add a configuration section to conf/processing.conf (or the equivalent JSON/YAML configuration) with the same name as your class in lowercase, and set enabled = false. The RunProcessing engine checks this flag before instantiating your module. You can also implement an on_demand option to allow conditional execution only when explicitly requested, which is useful for expensive operations like third-party API lookups.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →