How to Integrate External Configuration Extraction Frameworks with CAPEv2: A Complete Guide
CAPEv2 integrates external configuration extraction frameworks like RATDecoders and MalDuck by dynamically loading parser modules from designated directories during the processing phase, enabling automatic malware config extraction through a pluggable architecture defined in lib/cuckoo/common/load_extra_modules.py and cape_utils.py.
CAPEv2 supports seamless integration with third-party malware configuration extraction frameworks to enhance its static analysis capabilities. By leveraging a modular parser loading system, security analysts can integrate RATDecoders, MalDuck, DC3-MWCP, or MaCo directly into the sandbox processing pipeline. This guide explains how to configure and extend CAPEv2 to automatically extract command-and-control (C2) configurations and other malware indicators using these external tools.
Installation Requirements
Before enabling external parsers in CAPEv2, install the desired framework using pip within your CAPEv2 virtual environment. The frameworks must be accessible as Python packages to ensure the dynamic loader can import their modules correctly.
Install RATDecoders directly from the repository:
poetry run pip install git+https://github.com/kevthehermit/RATDecoders
For MalDuck, use the CERT-Polska repository:
poetry run pip install git+https://github.com/CERT-Polska/malduck
After installation, clone or copy the framework's parser modules into the CAPEv2 modules/processing/parsers/ directory structure. Each framework requires its own subdirectory containing valid Python packages with an extract_config(data) entry point.
Configuring processing.conf
Enable external frameworks in the CAPEv2 configuration file located at conf/default/processing.conf.default (lines 294-316). Create a dedicated section for each framework with the enabled flag and modules_path pointing to the parser directory.
Add the following configuration for RATDecoders:
[ratdecoders]
enabled = true
modules_path = modules/processing/parsers/RATDecoders/
# Optional: path to a custom YARA file that selects the samples
yara_path = modules/processing/parsers/RATDecoders/ratdecoders.yar
For MalDuck, the configuration follows an identical pattern, specifying the path to the MalDuck parser modules. CAPEv2 reads these settings at startup to determine which extraction frameworks to initialize during the analysis processing phase.
How CAPEv2 Loads External Parsers
The dynamic module loading mechanism resides in lib/cuckoo/common/load_extra_modules.py. This utility walks the specified parser directories using pkgutil.walk_packages and imports every Python file as a module under the modules.processing.parsers.<Framework> namespace.
For RATDecoders, the loader function (lines 40-45) performs the following operations:
# lib/cuckoo/common/load_extra_modules.py
def ratdecoders_load_modules(CUCKOO_ROOT: str):
parser_dir = os.path.join(CUCKOO_ROOT,
"modules", "processing", "parsers", "RATDecoders")
rat_modules = {}
for loader, mod_name, _ in pkgutil.walk_packages(
[parser_dir], "modules.processing.parsers.RATDecoders."):
try:
rat_modules[mod_name] = importlib.import_module(mod_name)
except Exception as e:
log.error("ratdecoder parser: %s – %s", mod_name, e)
return rat_modules
The MalDuck loader (lines 118-143) follows a similar pattern but returns specialized ExtractorModules objects that integrate with CAPEv2's YARA matching system. Both loaders execute during startup when cape_utils.py (lines 70-84) checks the configuration and initializes manager objects for enabled frameworks.
The Extraction Pipeline
Once loaded, external parsers execute during the process_extraction phase in lib/cuckoo/common/cape_utils.py. This function iterates over loaded parser collections and invokes each module's extract_config(file_bytes) method on samples matching the framework's YARA signatures.
The extraction workflow proceeds as follows:
- YARA Matching – CAPEv2 applies framework-specific YARA rules to submitted samples, filtering which files pass to each extractor.
- Config Extraction – Valid samples trigger calls to
extract_config(data), returning dictionaries of extracted key-value pairs. - Report Integration – Results merge into the analysis JSON under framework-specific keys like
ratdecodersormalduck.
The core extraction logic appears in cape_utils.py:
# lib/cuckoo/common/cape_utils.py
if process_cfg.ratdecoders.enabled:
rat_modules = load_extra_modules.ratdecoders_load_modules(CUCKOO_ROOT)
for name, mod in rat_modules.items():
cfg = mod.extract_config(file_bytes)
if cfg:
report["ratdecoders"][name] = cfg
Displaying Results in the Web UI
Extracted configurations surface in the CAPEv2 web interface through the Django template web/templates/analysis/static.html. The template renders framework-specific JSON data when present in the analysis report.
Access extracted data using the framework key defined during processing:
{% if report.ratdecoders %}
<h4>RATDecoders extraction</h4>
<pre>{{ report.ratdecoders|json_script:"ratdecoders-data" }}</pre>
{% endif %}
The web UI automatically displays configurations for any enabled framework following this pattern, requiring no additional template modifications when adding new parsers that follow the standard output format.
Summary
- Install frameworks using
pip install git+https://...before enabling them in CAPEv2. - Configure
processing.confto setenabled = trueand specify themodules_pathfor each external parser. - Dynamic loading occurs in
load_extra_modules.py, which walks parser directories and imports modules withpkgutil.walk_packages. - Extraction execution happens in
cape_utils.process_extraction(), callingextract_config(data)on matching samples. - Web display uses the analysis report keys (e.g.,
report.ratdecoders) rendered inweb/templates/analysis/static.html.
Frequently Asked Questions
What is the required entry point function for external parser modules?
External configuration extraction frameworks must implement def extract_config(data): as their primary entry point. This function receives the binary file bytes as input and returns a dictionary containing the extracted configuration keys and values. CAPEv2's process_extraction function in lib/cuckoo/common/cape_utils.py calls this method for every loaded parser module when processing matching samples.
How does CAPEv2 determine which samples to send to external parsers?
CAPEv2 uses YARA rule matching to filter samples before extraction. Each external framework can specify a yara_path in processing.conf pointing to a YARA rule file. The system applies these rules to submitted samples, and only files matching the signatures are passed to the corresponding extract_config functions. This prevents unnecessary processing of incompatible file types.
Can I integrate custom or proprietary configuration extractors with CAPEv2?
Yes. Any Python package implementing the extract_config(data) interface can integrate with CAPEv2. Create a directory under modules/processing/parsers/, add your modules with the required entry point, and add a configuration section to processing.conf with enabled = true and the appropriate modules_path. The dynamic loader in load_extra_modules.py will automatically discover and import your modules at startup.
Where are extracted configurations stored in the final analysis report?
Extracted configurations appear as top-level keys in the analysis JSON report, named according to the framework identifier (e.g., ratdecoders, malduck). These keys contain nested dictionaries mapping module names to their extracted configuration data. The web UI accesses these values through the report object (e.g., report.ratdecoders) and renders them in the Static Analysis tab.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →