# Storing and Retrieving CAPEv2 Analysis Results: MongoDB and Elasticsearch Options

> Explore MongoDB and Elasticsearch options for storing and retrieving CAPEv2 analysis results. Learn how to configure and utilize these powerful backends for your malware analysis workflow.

- Repository: [Kevin O'Reilly/capev2](https://github.com/kevoreilly/capev2)
- Tags: tutorial
- Published: 2026-03-05

---

**CAPEv2 supports two interchangeable backends for storing and retrieving analysis results—MongoDB and Elasticsearch—configured via [`reporting.conf`](https://github.com/kevoreilly/capev2/blob/main/reporting.conf) and abstracted through unified helper functions in [`dev_utils/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/mongodb.py) and [`dev_utils/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/elasticsearchdb.py).**

When analyzing malware with [CAPEv2](https://github.com/kevoreilly/capev2), the sandbox generates JSON-structured reports that must be persisted for later retrieval by the web interface, API, and internal utilities. The repository provides dual storage options for these CAPEv2 analysis results, allowing operators to choose between a document-oriented database or a scalable search engine without modifying application logic.

## Available Storage Backends for CAPEv2 Analysis Results

CAPEv2 implements a pluggable reporting architecture where both backends inherit from the generic `Report` abstract class. This design ensures that the web UI and API consume analysis data through identical helper functions regardless of which storage engine is active.

### MongoDB Document Store

The MongoDB implementation resides in [`modules/reporting/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/mongodb.py). This backend stores analysis reports as BSON documents in the **analysis** collection and maintains schema versioning in the **cuckoo_schema** collection. It handles BSON size limitations through a "loop saver" fallback mechanism for oversized documents.

### Elasticsearch Search Engine

The Elasticsearch implementation is located in [`modules/reporting/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/elasticsearchdb.py). This backend indexes reports into daily indices (e.g., `cuckoo-2024.09.01`) and provides advanced search capabilities. It supports a **search-only mode** that allows Elasticsearch to handle reads while MongoDB handles writes, or full read-write operation.

## Configuring MongoDB and Elasticsearch in CAPEv2

Storage backend selection occurs through [`reporting.conf`](https://github.com/kevoreilly/capev2/blob/main/reporting.conf), with validation enforced in the web application settings.

### Enabling MongoDB

To store and retrieve CAPEv2 analysis results using MongoDB, configure the `[mongodb]` section in [`conf/reporting.conf`](https://github.com/kevoreilly/capev2/blob/main/conf/reporting.conf):

```ini
[mongodb]
enabled = true
host = 127.0.0.1
port = 27017
db = cuckoo
username = <user>
password = <pass>

```

### Enabling Elasticsearch

For Elasticsearch storage, enable the `[elasticsearchdb]` section. Set `searchonly = true` to use Elasticsearch for reads while writing to MongoDB:

```ini
[elasticsearchdb]
enabled = true
searchonly = true
host = 127.0.0.1
port = 9200
index = cuckoo

```

### Validation in web/settings.py

The file [`web/web/settings.py`](https://github.com/kevoreilly/capev2/blob/main/web/web/settings.py) enforces that exactly one backend is active for write operations:

```python

# web/web/settings.py

if not cfg.mongodb.get("enabled") and not cfg.elasticsearchdb.get("enabled"):
    raise Exception("No database backend reporting module is enabled! Please enable either ElasticSearch or MongoDB.")
if cfg.mongodb.get("enabled") and cfg.elasticsearchdb.get("enabled") and not cfg.elasticsearchdb.get("searchonly"):
    raise Exception("Both database backend reporting modules are enabled. Please only enable ElasticSearch or MongoDB.")

```

## How CAPEv2 Writes Analysis Results

Both backends implement the `run()` method from the `Report` abstract class, processing the JSON-structured `results` dictionary generated by the analysis pipeline.

### MongoDB Write Path

In [`modules/reporting/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/mongodb.py), the `MongoDB.run()` method handles schema validation, call insertion, and document storage:

```python

# modules/reporting/mongodb.py (excerpt)

class MongoDB(Report):
    def run(self, results):
        if not HAVE_MONGO:
            raise CuckooDependencyError(...)
        # ensure schema version exists

        if "cuckoo_schema" in mongo_collection_names():
            if mongo_find_one("cuckoo_schema", {}, {"version": 1})["version"] != self.SCHEMA_VERSION:
                CuckooReportError(...)
        else:
            mongo_insert_one("cuckoo_schema", {"version": self.SCHEMA_VERSION})

        report = get_json_document(results, self.analysis_path)
        mongo_delete_data(int(report["info"]["id"]))          # clean old data

        new_processes = insert_calls(report, mongodb=True)   # store calls separately

        report["behavior"]["processes"] = new_processes
        ensure_valid_utf8(report)
        mongo_insert_one("analysis", report)                 # final insert

```

### Elasticsearch Write Path

In [`modules/reporting/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/elasticsearchdb.py), the `ElasticSearchDB.run()` method formats dates, normalizes fields, and indexes documents into daily indices:

```python

# modules/reporting/elasticsearchdb.py (excerpt)

class ElasticSearchDB(Report):
    def run(self, results):
        if not HAVE_ELASTICSEARCH:
            raise CuckooDependencyError(...)
        self.connect()
        self.check_analysis_index()
        self.check_calls_index()

        report = get_json_document(results, self.analysis_path)
        self.fix_fields(report)                 # normalise data for ES

        report = json.loads(json.dumps(report, default=str), object_hook=self.date_hook)
        new_processes = insert_calls(report, elastic_db=elastic_handler)
        report["behavior"]["processes"] = new_processes

        delete_analysis_and_related_calls(report["info"]["id"])  # clean old data

        self.format_dates(report)
        ensure_valid_utf8(report)
        self.index_report(report)               # ES index call

```

## Retrieving CAPEv2 Analysis Results

The retrieval layer abstracts the underlying storage engine through unified helper functions, allowing the web UI and API to query data without knowing which backend is active.

### Unified Helper Functions

The file [`dev_utils/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/mongodb.py) provides the primary interface for reading analysis data:

```python

# dev_utils/mongodb.py (excerpt)

def mongo_find_one(collection, query, projection=None):
    # Returns a single document from the specified collection

    ...

def mongo_find(collection, query, projection=None, sort=None, limit=0):
    # Returns multiple documents matching the query

    ...

def mongo_aggregate(collection, pipeline):
    # Performs aggregation operations

    ...

```

When Elasticsearch is enabled, [`dev_utils/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/elasticsearchdb.py) provides shim functions with identical signatures that forward queries to the ES client, returning Python dictionaries in the same format as MongoDB.

### Query Examples

Retrieve a specific analysis report by task ID using the unified interface:

```python

# utils/process.py (excerpt)

def get_report(task_id):
    from dev_utils.mongodb import mongo_find_one
    return mongo_find_one("analysis", {"info.id": int(task_id)}, {"_id": 0})

```

Query Elasticsearch directly for advanced search scenarios:

```python
from dev_utils.elasticsearchdb import elastic_handler

index = "cuckoo-2024.09.01"
query = {
    "query": {"match": {"info.id": 42}},
    "_source": ["info", "behavior"]
}
resp = elastic_handler.search(index=index, body=query)
print(resp["hits"]["hits"][0]["_source"])

```

## Summary

- **CAPEv2 supports two interchangeable backends** for storing and retrieving analysis results: MongoDB ([`modules/reporting/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/mongodb.py)) and Elasticsearch ([`modules/reporting/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/elasticsearchdb.py)).
- **Configuration is centralized** in [`reporting.conf`](https://github.com/kevoreilly/capev2/blob/main/reporting.conf), with validation in [`web/web/settings.py`](https://github.com/kevoreilly/capev2/blob/main/web/web/settings.py) ensuring exactly one write backend is active (or Elasticsearch in search-only mode alongside MongoDB).
- **Unified retrieval interface** provided by [`dev_utils/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/mongodb.py) and [`dev_utils/elasticsearchdb.py`](https://github.com/kevoreilly/capev2/blob/main/dev_utils/elasticsearchdb.py) allows the web UI, API, and utilities to query data without knowing which storage engine is active.
- **Schema versioning** is handled automatically in MongoDB via the `cuckoo_schema` collection, while Elasticsearch uses daily indices with normalized date fields.

## Frequently Asked Questions

### Can I run both MongoDB and Elasticsearch simultaneously?

Yes, but with limitations. You can enable both backends in [`reporting.conf`](https://github.com/kevoreilly/capev2/blob/main/reporting.conf) only if you set `searchonly = true` in the `[elasticsearchdb]` section. This configuration uses MongoDB as the primary write backend while allowing Elasticsearch to handle read operations for advanced search capabilities. The validation logic in [`web/web/settings.py`](https://github.com/kevoreilly/capev2/blob/main/web/web/settings.py) explicitly checks for this condition and raises an exception if both backends are enabled for writing.

### How does CAPEv2 handle large documents that exceed BSON limits?

The MongoDB reporting module in [`modules/reporting/mongodb.py`](https://github.com/kevoreilly/capev2/blob/main/modules/reporting/mongodb.py) includes a fallback mechanism for documents that exceed BSON's 16MB size limit. When processing behavior data, the `insert_calls()` function stores API call data separately, and the main report references these chunks. If a document still exceeds limits during the final `mongo_insert_one()` operation, the system can fall back to a "loop saver" approach that splits the data across multiple documents or collections while maintaining referential integrity through the analysis ID.

### What is the difference between normal and search-only Elasticsearch mode?

Normal Elasticsearch mode (`searchonly = false`) configures CAPEv2 to write analysis results directly to Elasticsearch indices, bypassing MongoDB entirely. In search-only mode (`searchonly = true`), CAPEv2 writes all data to MongoDB while maintaining a read-only Elasticsearch instance that indexes the same data for advanced querying. This hybrid approach is useful when you need MongoDB's reliability for primary storage but want Elasticsearch's full-text search capabilities for the web interface and API queries.

### Which backend is recommended for large-scale CAPEv2 deployments?

For large-scale deployments processing thousands of samples daily, **Elasticsearch** is generally recommended due to its horizontal scalability and superior full-text search performance across distributed clusters. However, if your deployment requires complex aggregation pipelines or you have strict consistency requirements for behavioral data, **MongoDB** may be preferable. Many production environments use the hybrid approach—MongoDB as the primary store with Elasticsearch in search-only mode—to leverage the strengths of both systems while maintaining data durability.