OpenCTI Retention Manager: How It Automates Data Lifecycle and Stale Data Removal
The OpenCTI retention manager is a background service that automatically purges stale data based on configurable retention rules, running every 30 seconds to evaluate filters and delete matching entities in batches of 1,500 while respecting concurrency limits.
The OpenCTI retention manager helps organizations maintain database hygiene by enforcing data retention policies on knowledge objects, files, and workbenches. According to the OpenCTI-Platform/opencti source code, this configurable scheduler operates as a singleton process that evaluates GraphQL-defined rules and performs batch deletions through Elasticsearch queries.
What Is the OpenCTI Retention Manager?
The OpenCTI retention manager is a periodic background worker implemented in src/manager/retentionManager.ts that enforces data lifecycle policies across the platform. It evaluates retention rules—user-defined configurations that specify what to keep and which entities to target—and automatically removes elements that exceed the specified age criteria.
The manager supports three distinct scopes exported as RETENTION_SCOPE_VALUES:
- Knowledge – STIX entities and relationships
- Files – Uploaded documents and imports
- Workbenches – Draft analysis workspaces
Each rule combines a max_retention value, a retention_unit (days, hours, etc.), and a set of filters to precisely target stale data.
How the Retention Manager Works
Initialization and Scheduling
The manager initializes when the platform starts via an import in src/manager/index.ts:
import './retentionManager';
The service runs continuously with a default interval of 30 seconds (RETENTION_MANAGER_INTERVAL), configurable via retention_manager:interval in app.yml or the RETENTION_MANAGER_INTERVAL environment variable.
Distributed Locking
To prevent race conditions in clustered deployments, the manager acquires a distributed lock using RETENTION_MANAGER_LOCK_KEY (default: retention_manager_lock). If another instance holds the lock, the current instance exits immediately. This guarantees that only one retention process runs at a time across the entire platform.
Rule Processing Flow
When the lock is acquired, the retentionHandler function executes the following logic:
- Fetch active rules – Calls
findRetentionRulesToExecute(context, RETENTION_MANAGER_USER)from the domain layer - Iterate rules – For each rule, invokes
executeProcessing(context, rule) - Calculate cutoff date – Computes the oldest allowed
updated_attimestamp usingmax_retentionandretention_unit - Find candidates – Executes
getElementsToDelete(context, rule)to build and run an Elasticsearch query combining the rule’sfiltersandscope
Batch Deletion Mechanics
Deletions proceed in controlled batches to prevent database overload:
- Batch size –
RETENTION_BATCH_SIZEdefaults to 1,500 elements per batch - Concurrency –
RETENTION_MAX_CONCURRENCYlimits parallel deletion workers to 2 concurrent operations - Deletion method – Elements are removed using
deleteElementwithin the batched loop
The manager logs each operation using logApp.debug and logApp.info for observability, then releases the lock upon completion.
Core Configuration Options
Configure the retention manager in app.yml or via environment variables:
| Setting | Default | Description |
|---|---|---|
retention_manager:enabled |
false |
Master switch to enable/disable the service |
retention_manager:interval |
30000 |
Execution frequency in milliseconds |
retention_manager:batch_size |
1500 |
Number of elements deleted per batch |
retention_manager:max_deletion_concurrency |
2 |
Maximum parallel deletion workers |
retention_manager:lock_key |
retention_manager_lock |
Distributed lock identifier |
Set retention_manager:enabled: true and restart the platform to activate the service.
Creating Retention Rules via GraphQL
Administrators define retention rules through the GraphQL API implemented in src/resolvers/retentionRule.js. The retentionRuleAdd mutation creates a new rule:
mutation CreateRule($input: RetentionRuleAddInput!) {
retentionRuleAdd(input: $input) {
id
name
max_retention
retention_unit
scope
}
}
Example variables for deleting knowledge older than 30 days:
{
"input": {
"name": "Knowledge older than 30 days",
"max_retention": 30,
"retention_unit": "days",
"scope": "knowledge",
"filters": [
{ "key": "updated_at", "operator": "lt", "value": "now-30d" }
]
}
}
Domain logic for rule validation and storage resides in src/domain/retentionRule.ts, which exports createRetentionRule, checkRetentionRule, and deleteRetentionRule.
Key Source Files and Implementation Details
| File | Purpose |
|---|---|
src/manager/retentionManager.ts |
Core scheduler, retentionHandler, batch deletion logic, and constants |
src/domain/retentionRule.ts |
Business logic for rule CRUD operations and validation |
src/resolvers/retentionRule.js |
GraphQL schema bindings for mutations and queries |
src/manager/index.ts |
Module loader that initializes the retention manager on startup |
src/tests/03-integration/04-manager/retentionManager-test.ts |
Integration tests for rule execution and element selection |
src/utils/platformModulesHelper.ts |
Exposes isRetentionManagerEnable() for UI components |
src/private/components/settings/Retention.tsx |
Frontend interface for rule management |
Summary
- The OpenCTI retention manager is a singleton background service that automatically deletes stale data according to user-defined retention rules.
- It processes rules every 30 seconds (configurable) using a distributed lock to prevent concurrent execution.
- Deletions occur in batches of 1,500 elements with a maximum concurrency of 2 parallel workers to ensure system stability.
- Rules target specific scopes (Knowledge, Files, Workbenches) and rely on Elasticsearch filters to identify candidates.
- Configuration and rule management are accessible via GraphQL mutations and the platform settings UI.
Frequently Asked Questions
How do I enable the retention manager in OpenCTI?
Set retention_manager:enabled: true in your app.yml configuration file or export RETENTION_MANAGER_ENABLED=true as an environment variable, then restart the platform. Verify activation by checking logs for retention manager entries or querying the isRetentionManagerEnable helper.
What is the default batch size for deletions?
The default batch size is 1,500 elements, controlled by the RETENTION_BATCH_SIZE constant in src/manager/retentionManager.ts. You can override this via the retention_manager:batch_size configuration setting to tune performance based on your Elasticsearch cluster capacity.
How does the retention manager prevent concurrent execution?
The manager uses a distributed lock with the key RETENTION_MANAGER_LOCK_KEY (default: retention_manager_lock). When the retentionHandler starts, it attempts to acquire this lock; if another instance holds it, the process exits immediately. This ensures only one node processes retention rules in a clustered deployment.
What entity types can retention rules target?
Retention rules support three scopes defined in RETENTION_SCOPE_VALUES: Knowledge (STIX entities and relationships), Files (uploaded documents), and Workbenches (draft analysis workspaces). Each rule specifies its target scope alongside filters to narrow the selection further.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →