Storing and Retrieving CAPEv2 Analysis Results: MongoDB and Elasticsearch Options

CAPEv2 supports two interchangeable backends for storing and retrieving analysis results—MongoDB and Elasticsearch—configured via reporting.conf and abstracted through unified helper functions in dev_utils/mongodb.py and dev_utils/elasticsearchdb.py.

When analyzing malware with CAPEv2, the sandbox generates JSON-structured reports that must be persisted for later retrieval by the web interface, API, and internal utilities. The repository provides dual storage options for these CAPEv2 analysis results, allowing operators to choose between a document-oriented database or a scalable search engine without modifying application logic.

Available Storage Backends for CAPEv2 Analysis Results

CAPEv2 implements a pluggable reporting architecture where both backends inherit from the generic Report abstract class. This design ensures that the web UI and API consume analysis data through identical helper functions regardless of which storage engine is active.

MongoDB Document Store

The MongoDB implementation resides in modules/reporting/mongodb.py. This backend stores analysis reports as BSON documents in the analysis collection and maintains schema versioning in the cuckoo_schema collection. It handles BSON size limitations through a "loop saver" fallback mechanism for oversized documents.

Elasticsearch Search Engine

The Elasticsearch implementation is located in modules/reporting/elasticsearchdb.py. This backend indexes reports into daily indices (e.g., cuckoo-2024.09.01) and provides advanced search capabilities. It supports a search-only mode that allows Elasticsearch to handle reads while MongoDB handles writes, or full read-write operation.

Configuring MongoDB and Elasticsearch in CAPEv2

Storage backend selection occurs through reporting.conf, with validation enforced in the web application settings.

Enabling MongoDB

To store and retrieve CAPEv2 analysis results using MongoDB, configure the [mongodb] section in conf/reporting.conf:

[mongodb]
enabled = true
host = 127.0.0.1
port = 27017
db = cuckoo
username = <user>
password = <pass>

Enabling Elasticsearch

For Elasticsearch storage, enable the [elasticsearchdb] section. Set searchonly = true to use Elasticsearch for reads while writing to MongoDB:

[elasticsearchdb]
enabled = true
searchonly = true
host = 127.0.0.1
port = 9200
index = cuckoo

Validation in web/settings.py

The file web/web/settings.py enforces that exactly one backend is active for write operations:


# web/web/settings.py

if not cfg.mongodb.get("enabled") and not cfg.elasticsearchdb.get("enabled"):
    raise Exception("No database backend reporting module is enabled! Please enable either ElasticSearch or MongoDB.")
if cfg.mongodb.get("enabled") and cfg.elasticsearchdb.get("enabled") and not cfg.elasticsearchdb.get("searchonly"):
    raise Exception("Both database backend reporting modules are enabled. Please only enable ElasticSearch or MongoDB.")

How CAPEv2 Writes Analysis Results

Both backends implement the run() method from the Report abstract class, processing the JSON-structured results dictionary generated by the analysis pipeline.

MongoDB Write Path

In modules/reporting/mongodb.py, the MongoDB.run() method handles schema validation, call insertion, and document storage:


# modules/reporting/mongodb.py (excerpt)

class MongoDB(Report):
    def run(self, results):
        if not HAVE_MONGO:
            raise CuckooDependencyError(...)
        # ensure schema version exists

        if "cuckoo_schema" in mongo_collection_names():
            if mongo_find_one("cuckoo_schema", {}, {"version": 1})["version"] != self.SCHEMA_VERSION:
                CuckooReportError(...)
        else:
            mongo_insert_one("cuckoo_schema", {"version": self.SCHEMA_VERSION})

        report = get_json_document(results, self.analysis_path)
        mongo_delete_data(int(report["info"]["id"]))          # clean old data

        new_processes = insert_calls(report, mongodb=True)   # store calls separately

        report["behavior"]["processes"] = new_processes
        ensure_valid_utf8(report)
        mongo_insert_one("analysis", report)                 # final insert

Elasticsearch Write Path

In modules/reporting/elasticsearchdb.py, the ElasticSearchDB.run() method formats dates, normalizes fields, and indexes documents into daily indices:


# modules/reporting/elasticsearchdb.py (excerpt)

class ElasticSearchDB(Report):
    def run(self, results):
        if not HAVE_ELASTICSEARCH:
            raise CuckooDependencyError(...)
        self.connect()
        self.check_analysis_index()
        self.check_calls_index()

        report = get_json_document(results, self.analysis_path)
        self.fix_fields(report)                 # normalise data for ES

        report = json.loads(json.dumps(report, default=str), object_hook=self.date_hook)
        new_processes = insert_calls(report, elastic_db=elastic_handler)
        report["behavior"]["processes"] = new_processes

        delete_analysis_and_related_calls(report["info"]["id"])  # clean old data

        self.format_dates(report)
        ensure_valid_utf8(report)
        self.index_report(report)               # ES index call

Retrieving CAPEv2 Analysis Results

The retrieval layer abstracts the underlying storage engine through unified helper functions, allowing the web UI and API to query data without knowing which backend is active.

Unified Helper Functions

The file dev_utils/mongodb.py provides the primary interface for reading analysis data:


# dev_utils/mongodb.py (excerpt)

def mongo_find_one(collection, query, projection=None):
    # Returns a single document from the specified collection

    ...

def mongo_find(collection, query, projection=None, sort=None, limit=0):
    # Returns multiple documents matching the query

    ...

def mongo_aggregate(collection, pipeline):
    # Performs aggregation operations

    ...

When Elasticsearch is enabled, dev_utils/elasticsearchdb.py provides shim functions with identical signatures that forward queries to the ES client, returning Python dictionaries in the same format as MongoDB.

Query Examples

Retrieve a specific analysis report by task ID using the unified interface:


# utils/process.py (excerpt)

def get_report(task_id):
    from dev_utils.mongodb import mongo_find_one
    return mongo_find_one("analysis", {"info.id": int(task_id)}, {"_id": 0})

Query Elasticsearch directly for advanced search scenarios:

from dev_utils.elasticsearchdb import elastic_handler

index = "cuckoo-2024.09.01"
query = {
    "query": {"match": {"info.id": 42}},
    "_source": ["info", "behavior"]
}
resp = elastic_handler.search(index=index, body=query)
print(resp["hits"]["hits"][0]["_source"])

Summary

  • CAPEv2 supports two interchangeable backends for storing and retrieving analysis results: MongoDB (modules/reporting/mongodb.py) and Elasticsearch (modules/reporting/elasticsearchdb.py).
  • Configuration is centralized in reporting.conf, with validation in web/web/settings.py ensuring exactly one write backend is active (or Elasticsearch in search-only mode alongside MongoDB).
  • Unified retrieval interface provided by dev_utils/mongodb.py and dev_utils/elasticsearchdb.py allows the web UI, API, and utilities to query data without knowing which storage engine is active.
  • Schema versioning is handled automatically in MongoDB via the cuckoo_schema collection, while Elasticsearch uses daily indices with normalized date fields.

Frequently Asked Questions

Can I run both MongoDB and Elasticsearch simultaneously?

Yes, but with limitations. You can enable both backends in reporting.conf only if you set searchonly = true in the [elasticsearchdb] section. This configuration uses MongoDB as the primary write backend while allowing Elasticsearch to handle read operations for advanced search capabilities. The validation logic in web/web/settings.py explicitly checks for this condition and raises an exception if both backends are enabled for writing.

How does CAPEv2 handle large documents that exceed BSON limits?

The MongoDB reporting module in modules/reporting/mongodb.py includes a fallback mechanism for documents that exceed BSON's 16MB size limit. When processing behavior data, the insert_calls() function stores API call data separately, and the main report references these chunks. If a document still exceeds limits during the final mongo_insert_one() operation, the system can fall back to a "loop saver" approach that splits the data across multiple documents or collections while maintaining referential integrity through the analysis ID.

What is the difference between normal and search-only Elasticsearch mode?

Normal Elasticsearch mode (searchonly = false) configures CAPEv2 to write analysis results directly to Elasticsearch indices, bypassing MongoDB entirely. In search-only mode (searchonly = true), CAPEv2 writes all data to MongoDB while maintaining a read-only Elasticsearch instance that indexes the same data for advanced querying. This hybrid approach is useful when you need MongoDB's reliability for primary storage but want Elasticsearch's full-text search capabilities for the web interface and API queries.

For large-scale deployments processing thousands of samples daily, Elasticsearch is generally recommended due to its horizontal scalability and superior full-text search performance across distributed clusters. However, if your deployment requires complex aggregation pipelines or you have strict consistency requirements for behavioral data, MongoDB may be preferable. Many production environments use the hybrid approach—MongoDB as the primary store with Elasticsearch in search-only mode—to leverage the strengths of both systems while maintaining data durability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →