OpenCTI Retention Manager: How It Automates Data Lifecycle and Stale Data Removal

The OpenCTI retention manager is a background service that automatically purges stale data based on configurable retention rules, running every 30 seconds to evaluate filters and delete matching entities in batches of 1,500 while respecting concurrency limits.

The OpenCTI retention manager helps organizations maintain database hygiene by enforcing data retention policies on knowledge objects, files, and workbenches. According to the OpenCTI-Platform/opencti source code, this configurable scheduler operates as a singleton process that evaluates GraphQL-defined rules and performs batch deletions through Elasticsearch queries.

What Is the OpenCTI Retention Manager?

The OpenCTI retention manager is a periodic background worker implemented in src/manager/retentionManager.ts that enforces data lifecycle policies across the platform. It evaluates retention rules—user-defined configurations that specify what to keep and which entities to target—and automatically removes elements that exceed the specified age criteria.

The manager supports three distinct scopes exported as RETENTION_SCOPE_VALUES:

  • Knowledge – STIX entities and relationships
  • Files – Uploaded documents and imports
  • Workbenches – Draft analysis workspaces

Each rule combines a max_retention value, a retention_unit (days, hours, etc.), and a set of filters to precisely target stale data.

How the Retention Manager Works

Initialization and Scheduling

The manager initializes when the platform starts via an import in src/manager/index.ts:

import './retentionManager';

The service runs continuously with a default interval of 30 seconds (RETENTION_MANAGER_INTERVAL), configurable via retention_manager:interval in app.yml or the RETENTION_MANAGER_INTERVAL environment variable.

Distributed Locking

To prevent race conditions in clustered deployments, the manager acquires a distributed lock using RETENTION_MANAGER_LOCK_KEY (default: retention_manager_lock). If another instance holds the lock, the current instance exits immediately. This guarantees that only one retention process runs at a time across the entire platform.

Rule Processing Flow

When the lock is acquired, the retentionHandler function executes the following logic:

  1. Fetch active rules – Calls findRetentionRulesToExecute(context, RETENTION_MANAGER_USER) from the domain layer
  2. Iterate rules – For each rule, invokes executeProcessing(context, rule)
  3. Calculate cutoff date – Computes the oldest allowed updated_at timestamp using max_retention and retention_unit
  4. Find candidates – Executes getElementsToDelete(context, rule) to build and run an Elasticsearch query combining the rule’s filters and scope

Batch Deletion Mechanics

Deletions proceed in controlled batches to prevent database overload:

  • Batch size – RETENTION_BATCH_SIZE defaults to 1,500 elements per batch
  • Concurrency – RETENTION_MAX_CONCURRENCY limits parallel deletion workers to 2 concurrent operations
  • Deletion method – Elements are removed using deleteElement within the batched loop

The manager logs each operation using logApp.debug and logApp.info for observability, then releases the lock upon completion.

Core Configuration Options

Configure the retention manager in app.yml or via environment variables:

Setting Default Description
retention_manager:enabled false Master switch to enable/disable the service
retention_manager:interval 30000 Execution frequency in milliseconds
retention_manager:batch_size 1500 Number of elements deleted per batch
retention_manager:max_deletion_concurrency 2 Maximum parallel deletion workers
retention_manager:lock_key retention_manager_lock Distributed lock identifier

Set retention_manager:enabled: true and restart the platform to activate the service.

Creating Retention Rules via GraphQL

Administrators define retention rules through the GraphQL API implemented in src/resolvers/retentionRule.js. The retentionRuleAdd mutation creates a new rule:

mutation CreateRule($input: RetentionRuleAddInput!) {
  retentionRuleAdd(input: $input) {
    id
    name
    max_retention
    retention_unit
    scope
  }
}

Example variables for deleting knowledge older than 30 days:

{
  "input": {
    "name": "Knowledge older than 30 days",
    "max_retention": 30,
    "retention_unit": "days",
    "scope": "knowledge",
    "filters": [
      { "key": "updated_at", "operator": "lt", "value": "now-30d" }
    ]
  }
}

Domain logic for rule validation and storage resides in src/domain/retentionRule.ts, which exports createRetentionRule, checkRetentionRule, and deleteRetentionRule.

Key Source Files and Implementation Details

File Purpose
src/manager/retentionManager.ts Core scheduler, retentionHandler, batch deletion logic, and constants
src/domain/retentionRule.ts Business logic for rule CRUD operations and validation
src/resolvers/retentionRule.js GraphQL schema bindings for mutations and queries
src/manager/index.ts Module loader that initializes the retention manager on startup
src/tests/03-integration/04-manager/retentionManager-test.ts Integration tests for rule execution and element selection
src/utils/platformModulesHelper.ts Exposes isRetentionManagerEnable() for UI components
src/private/components/settings/Retention.tsx Frontend interface for rule management

Summary

  • The OpenCTI retention manager is a singleton background service that automatically deletes stale data according to user-defined retention rules.
  • It processes rules every 30 seconds (configurable) using a distributed lock to prevent concurrent execution.
  • Deletions occur in batches of 1,500 elements with a maximum concurrency of 2 parallel workers to ensure system stability.
  • Rules target specific scopes (Knowledge, Files, Workbenches) and rely on Elasticsearch filters to identify candidates.
  • Configuration and rule management are accessible via GraphQL mutations and the platform settings UI.

Frequently Asked Questions

How do I enable the retention manager in OpenCTI?

Set retention_manager:enabled: true in your app.yml configuration file or export RETENTION_MANAGER_ENABLED=true as an environment variable, then restart the platform. Verify activation by checking logs for retention manager entries or querying the isRetentionManagerEnable helper.

What is the default batch size for deletions?

The default batch size is 1,500 elements, controlled by the RETENTION_BATCH_SIZE constant in src/manager/retentionManager.ts. You can override this via the retention_manager:batch_size configuration setting to tune performance based on your Elasticsearch cluster capacity.

How does the retention manager prevent concurrent execution?

The manager uses a distributed lock with the key RETENTION_MANAGER_LOCK_KEY (default: retention_manager_lock). When the retentionHandler starts, it attempts to acquire this lock; if another instance holds it, the process exits immediately. This ensures only one node processes retention rules in a clustered deployment.

What entity types can retention rules target?

Retention rules support three scopes defined in RETENTION_SCOPE_VALUES: Knowledge (STIX entities and relationships), Files (uploaded documents), and Workbenches (draft analysis workspaces). Each rule specifies its target scope alongside filters to narrow the selection further.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →