Storing and Retrieving CAPEv2 Analysis Results: MongoDB and Elasticsearch Options
CAPEv2 supports two interchangeable backends for storing and retrieving analysis results—MongoDB and Elasticsearch—configured via reporting.conf and abstracted through unified helper functions in dev_utils/mongodb.py and dev_utils/elasticsearchdb.py.
When analyzing malware with CAPEv2, the sandbox generates JSON-structured reports that must be persisted for later retrieval by the web interface, API, and internal utilities. The repository provides dual storage options for these CAPEv2 analysis results, allowing operators to choose between a document-oriented database or a scalable search engine without modifying application logic.
Available Storage Backends for CAPEv2 Analysis Results
CAPEv2 implements a pluggable reporting architecture where both backends inherit from the generic Report abstract class. This design ensures that the web UI and API consume analysis data through identical helper functions regardless of which storage engine is active.
MongoDB Document Store
The MongoDB implementation resides in modules/reporting/mongodb.py. This backend stores analysis reports as BSON documents in the analysis collection and maintains schema versioning in the cuckoo_schema collection. It handles BSON size limitations through a "loop saver" fallback mechanism for oversized documents.
Elasticsearch Search Engine
The Elasticsearch implementation is located in modules/reporting/elasticsearchdb.py. This backend indexes reports into daily indices (e.g., cuckoo-2024.09.01) and provides advanced search capabilities. It supports a search-only mode that allows Elasticsearch to handle reads while MongoDB handles writes, or full read-write operation.
Configuring MongoDB and Elasticsearch in CAPEv2
Storage backend selection occurs through reporting.conf, with validation enforced in the web application settings.
Enabling MongoDB
To store and retrieve CAPEv2 analysis results using MongoDB, configure the [mongodb] section in conf/reporting.conf:
[mongodb]
enabled = true
host = 127.0.0.1
port = 27017
db = cuckoo
username = <user>
password = <pass>
Enabling Elasticsearch
For Elasticsearch storage, enable the [elasticsearchdb] section. Set searchonly = true to use Elasticsearch for reads while writing to MongoDB:
[elasticsearchdb]
enabled = true
searchonly = true
host = 127.0.0.1
port = 9200
index = cuckoo
Validation in web/settings.py
The file web/web/settings.py enforces that exactly one backend is active for write operations:
# web/web/settings.py
if not cfg.mongodb.get("enabled") and not cfg.elasticsearchdb.get("enabled"):
raise Exception("No database backend reporting module is enabled! Please enable either ElasticSearch or MongoDB.")
if cfg.mongodb.get("enabled") and cfg.elasticsearchdb.get("enabled") and not cfg.elasticsearchdb.get("searchonly"):
raise Exception("Both database backend reporting modules are enabled. Please only enable ElasticSearch or MongoDB.")
How CAPEv2 Writes Analysis Results
Both backends implement the run() method from the Report abstract class, processing the JSON-structured results dictionary generated by the analysis pipeline.
MongoDB Write Path
In modules/reporting/mongodb.py, the MongoDB.run() method handles schema validation, call insertion, and document storage:
# modules/reporting/mongodb.py (excerpt)
class MongoDB(Report):
def run(self, results):
if not HAVE_MONGO:
raise CuckooDependencyError(...)
# ensure schema version exists
if "cuckoo_schema" in mongo_collection_names():
if mongo_find_one("cuckoo_schema", {}, {"version": 1})["version"] != self.SCHEMA_VERSION:
CuckooReportError(...)
else:
mongo_insert_one("cuckoo_schema", {"version": self.SCHEMA_VERSION})
report = get_json_document(results, self.analysis_path)
mongo_delete_data(int(report["info"]["id"])) # clean old data
new_processes = insert_calls(report, mongodb=True) # store calls separately
report["behavior"]["processes"] = new_processes
ensure_valid_utf8(report)
mongo_insert_one("analysis", report) # final insert
Elasticsearch Write Path
In modules/reporting/elasticsearchdb.py, the ElasticSearchDB.run() method formats dates, normalizes fields, and indexes documents into daily indices:
# modules/reporting/elasticsearchdb.py (excerpt)
class ElasticSearchDB(Report):
def run(self, results):
if not HAVE_ELASTICSEARCH:
raise CuckooDependencyError(...)
self.connect()
self.check_analysis_index()
self.check_calls_index()
report = get_json_document(results, self.analysis_path)
self.fix_fields(report) # normalise data for ES
report = json.loads(json.dumps(report, default=str), object_hook=self.date_hook)
new_processes = insert_calls(report, elastic_db=elastic_handler)
report["behavior"]["processes"] = new_processes
delete_analysis_and_related_calls(report["info"]["id"]) # clean old data
self.format_dates(report)
ensure_valid_utf8(report)
self.index_report(report) # ES index call
Retrieving CAPEv2 Analysis Results
The retrieval layer abstracts the underlying storage engine through unified helper functions, allowing the web UI and API to query data without knowing which backend is active.
Unified Helper Functions
The file dev_utils/mongodb.py provides the primary interface for reading analysis data:
# dev_utils/mongodb.py (excerpt)
def mongo_find_one(collection, query, projection=None):
# Returns a single document from the specified collection
...
def mongo_find(collection, query, projection=None, sort=None, limit=0):
# Returns multiple documents matching the query
...
def mongo_aggregate(collection, pipeline):
# Performs aggregation operations
...
When Elasticsearch is enabled, dev_utils/elasticsearchdb.py provides shim functions with identical signatures that forward queries to the ES client, returning Python dictionaries in the same format as MongoDB.
Query Examples
Retrieve a specific analysis report by task ID using the unified interface:
# utils/process.py (excerpt)
def get_report(task_id):
from dev_utils.mongodb import mongo_find_one
return mongo_find_one("analysis", {"info.id": int(task_id)}, {"_id": 0})
Query Elasticsearch directly for advanced search scenarios:
from dev_utils.elasticsearchdb import elastic_handler
index = "cuckoo-2024.09.01"
query = {
"query": {"match": {"info.id": 42}},
"_source": ["info", "behavior"]
}
resp = elastic_handler.search(index=index, body=query)
print(resp["hits"]["hits"][0]["_source"])
Summary
- CAPEv2 supports two interchangeable backends for storing and retrieving analysis results: MongoDB (
modules/reporting/mongodb.py) and Elasticsearch (modules/reporting/elasticsearchdb.py). - Configuration is centralized in
reporting.conf, with validation inweb/web/settings.pyensuring exactly one write backend is active (or Elasticsearch in search-only mode alongside MongoDB). - Unified retrieval interface provided by
dev_utils/mongodb.pyanddev_utils/elasticsearchdb.pyallows the web UI, API, and utilities to query data without knowing which storage engine is active. - Schema versioning is handled automatically in MongoDB via the
cuckoo_schemacollection, while Elasticsearch uses daily indices with normalized date fields.
Frequently Asked Questions
Can I run both MongoDB and Elasticsearch simultaneously?
Yes, but with limitations. You can enable both backends in reporting.conf only if you set searchonly = true in the [elasticsearchdb] section. This configuration uses MongoDB as the primary write backend while allowing Elasticsearch to handle read operations for advanced search capabilities. The validation logic in web/web/settings.py explicitly checks for this condition and raises an exception if both backends are enabled for writing.
How does CAPEv2 handle large documents that exceed BSON limits?
The MongoDB reporting module in modules/reporting/mongodb.py includes a fallback mechanism for documents that exceed BSON's 16MB size limit. When processing behavior data, the insert_calls() function stores API call data separately, and the main report references these chunks. If a document still exceeds limits during the final mongo_insert_one() operation, the system can fall back to a "loop saver" approach that splits the data across multiple documents or collections while maintaining referential integrity through the analysis ID.
What is the difference between normal and search-only Elasticsearch mode?
Normal Elasticsearch mode (searchonly = false) configures CAPEv2 to write analysis results directly to Elasticsearch indices, bypassing MongoDB entirely. In search-only mode (searchonly = true), CAPEv2 writes all data to MongoDB while maintaining a read-only Elasticsearch instance that indexes the same data for advanced querying. This hybrid approach is useful when you need MongoDB's reliability for primary storage but want Elasticsearch's full-text search capabilities for the web interface and API queries.
Which backend is recommended for large-scale CAPEv2 deployments?
For large-scale deployments processing thousands of samples daily, Elasticsearch is generally recommended due to its horizontal scalability and superior full-text search performance across distributed clusters. However, if your deployment requires complex aggregation pipelines or you have strict consistency requirements for behavioral data, MongoDB may be preferable. Many production environments use the hybrid approach—MongoDB as the primary store with Elasticsearch in search-only mode—to leverage the strengths of both systems while maintaining data durability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →