Choosing Between DynamoDB and Elasticsearch for AWS Data Lakes: An Architectural Guide

When architecting a data lake on Amazon Web Services, use DynamoDB for high-velocity, key-value ingestion and point lookups, while leveraging the Amazon Web Services Elasticsearch search engine for data lakes to power full-text search, aggregations, and complex analytics.

Building a scalable data lake on AWS often requires balancing high-throughput ingestion with rich query capabilities. While DynamoDB excels at handling massive write loads and millisecond-latency key lookups, the Amazon Web Services Elasticsearch search engine for data lakes provides the full-text search, aggregation framework, and semantic query support necessary for modern analytics workloads. This guide examines the architectural trade-offs between these services, referencing implementation details from the elastic/elasticsearch repository to demonstrate production-grade integration patterns.

Data Model and Query Capabilities

DynamoDB: Key-Value and Document Patterns

DynamoDB organizes data into tables with a mandatory primary key (partition key plus optional sort key). While it supports flexible document schemas, every access pattern must align with this key structure. The service provides simple point lookups, range queries on sort keys, and limited secondary indexes, but lacks native full-text search or complex aggregation capabilities.

Elasticsearch: Full-Text Search and Aggregations

Elasticsearch stores JSON documents without primary-key constraints, allowing any field to be indexed and queried using the rich Query DSL. According to the repository documentation, the engine supports k-NN semantic search, hybrid queries, and complex aggregations essential for data lake analytics. The query DSL examples for semantic and k-NN queries are documented in query-dsl-semantic-query.md and query-dsl-knn-query.md within the repository.

Performance and Cost Considerations

Latency and Throughput Patterns

DynamoDB delivers single-digit millisecond latency for point reads and writes, scaling horizontally through partition key distribution. It handles billions of writes daily with predictable throughput, making it ideal for high-velocity data lake ingestion.

Elasticsearch provides millisecond-level search latency, but indexing throughput depends on shard sizing and cluster resources. The repository's cloud-aws-best-practices.md (lines 12-30) emphasizes that scaling requires careful shard allocation and multi-AZ placement to maintain performance.

Cost Models for Data Lake Workloads

DynamoDB charges per request (RCU/WCU) or provisioned capacity, with no query-related storage charges. This model favors high-volume, predictable access patterns.

Elasticsearch pricing depends on node instances (CPU, RAM, storage) and data transfer costs. Storage costs are higher, particularly with replicated shards. For cost optimization, store raw immutable records in DynamoDB while maintaining a smaller, purpose-built search index in the Amazon Web Services Elasticsearch search engine for data lakes.

Operational and Security Architecture

AWS Deployment Best Practices

The repository's cloud-aws-best-practices.md (lines 12-30) provides specific guidance for AWS deployments. Key recommendations include using instance store for high-throughput data nodes, multi-AZ shard allocation for resilience, and provisioned IOPS EBS for master-eligible nodes. These settings ensure the cluster can handle data lake query loads while maintaining availability.

IAM and Connector Permissions

Securing the pipeline requires careful IAM configuration. DynamoDB needs stream read permissions for change data capture, while Elasticsearch requires specific connector permissions. According to ElasticServiceAccounts.java (lines 49-58), service accounts must possess cluster:admin/xpack/connector/* permissions to execute Elastic Connectors. This permission set enables the connector to manage sync jobs and write to target indices securely.

Integration Patterns: Connecting DynamoDB to Elasticsearch

For data lakes requiring both high-throughput storage and rich search capabilities, the recommended pattern stores raw events in DynamoDB while replicating searchable fields to Elasticsearch. The repository supports this through Elastic Connectors and custom pipeline implementations.

Java Bulk Indexing Implementation

When building custom sync pipelines, use the Bulk API for efficient document ingestion:

// DynamoDB client (v2)
DynamoDbClient ddb = DynamoDbClient.builder()
    .region(Region.US_EAST_1)
    .build();

// Scan the table (simple example – production should use pagination / parallel scans)
ScanResponse scan = ddb.scan(ScanRequest.builder()
    .tableName("my-lake-table")
    .build());

// Elasticsearch bulk request
BulkRequest.Builder bulk = new BulkRequest.Builder();

for (Map<String, AttributeValue> item : scan.items()) {
    // Convert DynamoDB map to JSON (using Jackson or manual mapping)
    String json = mapToJson(item);

    bulk.operations(op -> op
        .create(c -> c
            .index(i -> i
                .index("lake-index")
                .id(item.get("PK").s())
            )
            .document(json)
        )
    );
}

// Execute bulk request
ElasticsearchClient esClient = new ElasticsearchClient(
    new RestClientBuilder(HttpHost.create("https://my-es-domain.region.es.amazonaws.com"))
);
BulkResponse response = esClient.bulk(bulk.build());

// Handle errors
if (response.errors()) {
    response.items().stream()
        .filter(i -> i.error() != null)
        .forEach(i -> System.err.println("Failed: " + i.error().reason()));
}

The ElasticServiceAccounts.java file (lines 49-58) confirms that executing this bulk operation requires cluster:admin/xpack/connector/* permissions for connector-based service accounts.

Elastic Connector Configuration

For managed synchronization, configure the AWS DynamoDB source connector. The sync-job lifecycle is documented in es-sync-rules.md (line 67):


# file: dynamodb-connector.yml

service_type: aws-dynamodb
configuration:
  table: my-lake-table
  region: us-east-1
  access_key: ${AWS_ACCESS_KEY}
  secret_key: ${AWS_SECRET_KEY}
  scan_batch_size: 5000
  stream_enabled: true          # use DynamoDB Streams for incremental sync

  sync:
    schedule: "0 */5 * * * *"   # every 5 minutes

target_index:
  name: lake-index
  refresh: true
  mapping:
    properties:
      PK: { type: keyword }
      timestamp: { type: date }
      payload: { type: text }

Terraform Deployment

Provision the Amazon Web Services Elasticsearch search engine for data lakes with optimized settings from cloud-aws-best-practices.md (lines 12-30):

resource "aws_elasticsearch_domain" "lake_search" {
  domain_name = "lake-search"

  elasticsearch_version = "7.17"

  cluster_config {
    instance_type = "r5.large.elasticsearch"
    instance_count = 3
    zone_awareness_enabled = true
    zone_awareness_config {
      availability_zone_count = 2
    }
    // Use instance store for data nodes (recommended in cloud-aws-best-practices)
    storage_type = "instance"
  }

  ebs_options {
    ebs_enabled = false
  }

  zone_awareness_enabled = true

  advanced_options = {
    "rest.action.multi.allow_explicit_index" = "true"
  }

  // Enable fine-grained access control (required for connector IAM)
  advanced_security_options {
    enabled = true
    internal_user_database_enabled = true
  }
}

Summary

  • DynamoDB excels as a high-throughput, low-latency key-value store for raw data lake ingestion, offering predictable costs per request and automatic scaling.
  • Elasticsearch (Amazon OpenSearch Service) provides the full-text search, aggregations, and semantic query capabilities required for analytics workloads, as implemented in the elastic/elasticsearch repository.
  • Hybrid architectures leverage DynamoDB Streams or Elastic Connectors (documented in es-sync-rules.md) to maintain searchable indexes without duplicating storage costs.
  • AWS-specific optimizations include using instance storage for data nodes and multi-AZ deployment, detailed in cloud-aws-best-practices.md (lines 12-30).
  • Security integration requires granting cluster:admin/xpack/connector/* permissions to service accounts, as defined in ElasticServiceAccounts.java (lines 49-58).

Frequently Asked Questions

Can I use DynamoDB as the primary storage and Elasticsearch only for search indexes?

Yes, this is the recommended pattern for cost-efficient data lakes. Store immutable raw events in DynamoDB for cheap, scalable persistence, then replicate only the fields needed for search and analytics to Elasticsearch using DynamoDB Streams with Lambda or the Elastic Connector for DynamoDB. The connector's sync-job configuration is documented in es-sync-rules.md (line 67), allowing you to schedule incremental syncs every few minutes to keep indexes fresh without overloading your cluster.

What permissions does Elasticsearch require to sync data from DynamoDB?

Elasticsearch service accounts need the cluster:admin/xpack/connector/* permission set to manage connector jobs and write to target indices. According to ElasticServiceAccounts.java (lines 49-58), these permissions enable the connector to execute sync rules, handle bulk indexing, and manage pipeline configurations. Additionally, the AWS IAM role running the connector requires dynamodb:Scan, dynamodb:DescribeTable, and dynamodb:GetRecords permissions for table access and stream consumption.

How do I optimize Elasticsearch deployment costs on AWS for data lake workloads?

Deploy Elasticsearch with instance storage rather than EBS for data nodes to reduce storage costs while maintaining high throughput, as recommended in cloud-aws-best-practices.md (lines 12-30). Use multi-AZ deployments only for master-eligible nodes and critical data, and implement index lifecycle management to move older data to cheaper storage tiers. For data lakes specifically, maintain a smaller, curated search index in Elasticsearch while keeping raw historical data in DynamoDB or S3, minimizing the node count and instance sizes needed for your Elasticsearch cluster.

What are the latency differences between querying DynamoDB versus Elasticsearch?

DynamoDB provides single-digit millisecond latency for point reads and writes when accessing items by primary key, making it ideal for real-time ingestion and lookup patterns. Elasticsearch delivers millisecond-level search latency for complex queries, but this depends on shard sizing, query complexity, and cluster load. For data lake architectures, use DynamoDB when you need immediate retrieval of specific records by ID, and leverage Elasticsearch when you require ad-hoc search, aggregations, or fuzzy matching across large datasets, as supported by the Query DSL and k-NN capabilities in the elastic/elasticsearch repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →