How to Develop and Implement Custom Processing Modules for CAPEv2
Custom CAPEv2 processing modules are Python classes inheriting from the Processing base class that transform raw analysis artifacts into structured data by implementing a run() method and placing the file in modules/processing/.
CAPEv2 uses a modular pipeline to analyze malware samples, and understanding how to develop and implement custom processing modules for CAPEv2 allows you to extend its capabilities beyond the default behavior analysis. The framework automatically discovers and executes these modules, merging their output into the final analysis report without requiring modifications to the core scheduler.
Understanding the CAPEv2 Processing Pipeline Architecture
The Processing Base Class
Every processing module must inherit from the abstract Processing class defined in lib/cuckoo/common/abstracts.py. This base class provides the essential API contract that the orchestration engine expects, including helper methods for path resolution, task metadata access, and statistics tracking.
The Processing class defines three critical setter methods that the framework calls before run():
set_path()– Injects the absolute path to the analysis directoryset_task()– Provides the task dictionary containing target file, platform, and optionsset_options()– Supplies module-specific configuration fromprocessing.conf
How RunProcessing Orchestrates Modules
The RunProcessing class in lib/cuckoo/core/plugins.py manages the execution lifecycle. When a task completes, the scheduler instantiates RunProcessing and invokes its process() method. This method iterates over every Python file in modules/processing/, imports classes inheriting from Processing, checks the configuration for an enabled flag, and executes the run() method.
Timing data is automatically captured and stored in results["statistics"]["processing"] without requiring explicit instrumentation in your module.
Step-by-Step Guide to Creating Custom Processing Modules
Step 1: Create the Module File in modules/processing/
Navigate to the modules/processing/ directory and create a new Python file. The filename should be descriptive but does not need to match the class name. For example, pe_counter.py or network_enhancer.py.
Step 2: Inherit from the Processing Base Class
Import the base class and define your module:
from lib.cuckoo.common.abstracts import Processing
class PeCounter(Processing):
"""Counts PE files extracted during analysis."""
pass
Step 3: Implement the run() Method
The run() method performs your analysis logic and must return a dictionary. Set self.key to define where the data appears in the final report:
def run(self):
self.key = "pe_counter"
# Access analysis path via self.analysis_path
# Access task info via self.task
# Access config options via self.options
return {"count": 0, "files": []}
Step 4: Configure the Module in processing.conf
Add a section to conf/processing.conf (or the equivalent JSON/YAML configuration):
[pe_counter]
enabled = true
on_demand = false
The section name must match your class name in lowercase. Set enabled = true to activate the module.
Practical Code Examples for CAPEv2 Processing Modules
Example 1: Counting PE Files in Analysis Artifacts
This complete module demonstrates accessing the analysis directory and returning structured data:
# modules/processing/pe_counter.py
import os
from lib.cuckoo.common.abstracts import Processing
class PeCounter(Processing):
"""Counts PE files in the `files` directory of an analysis."""
def run(self):
# Define the key for results aggregation
self.key = "pe_counter"
# Construct path to extracted files
files_dir = os.path.join(self.analysis_path, "files")
# Safely handle missing directories
if not os.path.exists(files_dir):
return {"total_pe": 0, "pe_list": []}
pe_files = [f for f in os.listdir(files_dir) if f.lower().endswith(".exe")]
return {"total_pe": len(pe_files), "pe_list": pe_files}
The self.analysis_path attribute is automatically populated by the framework via set_path().
Example 2: Accessing Task Metadata and Previous Results
This module demonstrates reading from self.task and self.results to build composite reports:
# modules/processing/report_summary.py
from lib.cuckoo.common.abstracts import Processing
class ReportSummary(Processing):
"""Creates a summary using data from earlier processing modules."""
def run(self):
self.key = "summary"
# Access task metadata
target = self.task.get("target", "unknown")
platform = self.task.get("platform", "windows")
# Access results from other modules
behavior = self.results.get("behavior", {})
processes = behavior.get("processes", [])
num_procs = len(processes)
# Access memory module results if available
memory_data = self.results.get("memory", {})
mem_usage = memory_data.get("mem_usage", [])
avg_mem = sum(mem_usage) / len(mem_usage) if mem_usage else 0
return {
"target": target,
"platform": platform,
"process_count": num_procs,
"average_memory_mb": round(avg_mem, 2),
"generated_by": __class__.__name__
}
Example 3: Debugging Modules with RunProcessing
Test your module without running a full analysis using the RunProcessing class directly:
>>> from lib.cuckoo.core.plugins import RunProcessing
>>>
>>> # Create mock task and results structure
>>> task = {"id": 42, "platform": "windows", "target": "sample.exe"}
>>> results = {"statistics": {"processing": []}}
>>>
>>> # Initialize the runner
>>> runner = RunProcessing(task, results)
>>>
>>> # Import and test your custom module
>>> from modules.processing.pe_counter import PeCounter
>>> runner.process(PeCounter)
>>>
>>> # Inspect the results
>>> results['pe_counter']
{'total_pe': 3, 'pe_list': ['a.exe', 'b.exe', 'c.exe']}
This approach allows rapid iteration without submitting samples to the sandbox.
Advanced Configuration and Testing Strategies
Enabling On-Demand Processing
CAPEv2 supports conditional execution via the on_demand configuration flag. When set to true, the module only executes when explicitly requested rather than during every analysis:
[pe_counter]
enabled = true
on_demand = true
In your module, check this flag via self.options.get("on_demand") to implement conditional logic. This pattern is used by the url_analysis module to avoid unnecessary API calls to VirusTotal unless specifically enabled.
Unit Testing with ProcessingMock
The repository includes testing utilities in tests/processor_tests.py. Use the ProcessingMock pattern to instantiate your module with controlled inputs:
# tests/test_pe_counter.py
import unittest
from modules.processing.pe_counter import PeCounter
from tests.processor_tests import ProcessingMock
class TestPeCounter(unittest.TestCase):
def test_empty_directory(self):
# Setup mock with temporary path
mock = ProcessingMock()
module = PeCounter()
module.set_path("/tmp/empty_analysis")
module.set_task({"id": 1, "target": "test.exe"})
module.set_options({"enabled": "true"})
result = module.run()
self.assertEqual(result["total_pe"], 0)
This approach isolates your logic from the CAPEv2 runtime while verifying correct behavior.
Summary
- Inherit from
Processing: All custom modules must subclass the abstract base class inlib/cuckoo/common/abstracts.pyand implement therun()method. - Place files in
modules/processing/: TheRunProcessingengine automatically discovers Python files in this directory; no manual registration is required. - Return dictionaries: The
run()method must return a dict that gets merged into the global results under the key specified byself.key. - Configure via
processing.conf: Enable modules by adding a lowercase section matching the class name withenabled = true. - Leverage built-in helpers: Use
self.analysis_path,self.task, andself.optionsto access analysis data, and rely on automatic timing statistics collection.
Frequently Asked Questions
What is the minimum required method to implement in a CAPEv2 processing module?
You must implement the run() method. This method receives no arguments (all context is available via self attributes set by the framework) and must return a dictionary containing your analysis results. The base class Processing in lib/cuckoo/common/abstracts.py defines this as an abstract method, so the framework will raise an error if you omit it.
Where does CAPEv2 look for custom processing modules?
CAPEv2 scans the modules/processing/ directory for Python files containing classes that inherit from Processing. The RunProcessing class in lib/cuckoo/core/plugins.py handles this discovery automatically during initialization. You do not need to import or register your module manually; simply placing the file in the correct directory with the proper inheritance pattern makes it available to the framework.
How do I access analysis results from other processing modules in my custom module?
Access the global results dictionary via self.results, which the framework populates with data from previously executed modules. For example, to read behavioral data from the behavior module, use self.results.get("behavior", {}). To access the current task's metadata (such as target file or platform), use self.task. These attributes are injected by RunProcessing before your run() method executes.
Can I disable a processing module without removing the code?
Yes. Add a configuration section to conf/processing.conf (or the equivalent JSON/YAML configuration) with the same name as your class in lowercase, and set enabled = false. The RunProcessing engine checks this flag before instantiating your module. You can also implement an on_demand option to allow conditional execution only when explicitly requested, which is useful for expensive operations like third-party API lookups.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →