# Choosing Between DynamoDB and Elasticsearch for AWS Data Lakes: An Architectural Guide

> Architecting a data lake on AWS? Learn when to use DynamoDB for fast key-value ingestion and Elasticsearch for powerful search and analytics. Optimize your AWS Elasticsearch data lake.

- Repository: [elastic/elasticsearch](https://github.com/elastic/elasticsearch)
- Tags: architecture
- Published: 2026-02-16

---

**When architecting a data lake on Amazon Web Services, use DynamoDB for high-velocity, key-value ingestion and point lookups, while leveraging the Amazon Web Services Elasticsearch search engine for data lakes to power full-text search, aggregations, and complex analytics.**

Building a scalable data lake on AWS often requires balancing high-throughput ingestion with rich query capabilities. While DynamoDB excels at handling massive write loads and millisecond-latency key lookups, the Amazon Web Services Elasticsearch search engine for data lakes provides the full-text search, aggregation framework, and semantic query support necessary for modern analytics workloads. This guide examines the architectural trade-offs between these services, referencing implementation details from the `elastic/elasticsearch` repository to demonstrate production-grade integration patterns.

## Data Model and Query Capabilities

### DynamoDB: Key-Value and Document Patterns

DynamoDB organizes data into tables with a mandatory primary key (partition key plus optional sort key). While it supports flexible document schemas, every access pattern must align with this key structure. The service provides simple point lookups, range queries on sort keys, and limited secondary indexes, but lacks native full-text search or complex aggregation capabilities.

### Elasticsearch: Full-Text Search and Aggregations

Elasticsearch stores JSON documents without primary-key constraints, allowing any field to be indexed and queried using the rich Query DSL. According to the repository documentation, the engine supports **k-NN** semantic search, hybrid queries, and complex aggregations essential for data lake analytics. The query DSL examples for semantic and k-NN queries are documented in [`query-dsl-semantic-query.md`](https://github.com/elastic/elasticsearch/blob/main/query-dsl-semantic-query.md) and [`query-dsl-knn-query.md`](https://github.com/elastic/elasticsearch/blob/main/query-dsl-knn-query.md) within the repository.

## Performance and Cost Considerations

### Latency and Throughput Patterns

DynamoDB delivers single-digit millisecond latency for point reads and writes, scaling horizontally through partition key distribution. It handles billions of writes daily with predictable throughput, making it ideal for high-velocity data lake ingestion.

Elasticsearch provides millisecond-level search latency, but indexing throughput depends on shard sizing and cluster resources. The repository's [`cloud-aws-best-practices.md`](https://github.com/elastic/elasticsearch/blob/main/cloud-aws-best-practices.md) (lines 12-30) emphasizes that scaling requires careful shard allocation and multi-AZ placement to maintain performance.

### Cost Models for Data Lake Workloads

DynamoDB charges per request (RCU/WCU) or provisioned capacity, with no query-related storage charges. This model favors high-volume, predictable access patterns.

Elasticsearch pricing depends on node instances (CPU, RAM, storage) and data transfer costs. Storage costs are higher, particularly with replicated shards. For cost optimization, store raw immutable records in DynamoDB while maintaining a smaller, purpose-built search index in the Amazon Web Services Elasticsearch search engine for data lakes.

## Operational and Security Architecture

### AWS Deployment Best Practices

The repository's [`cloud-aws-best-practices.md`](https://github.com/elastic/elasticsearch/blob/main/cloud-aws-best-practices.md) (lines 12-30) provides specific guidance for AWS deployments. Key recommendations include using instance store for high-throughput data nodes, multi-AZ shard allocation for resilience, and provisioned IOPS EBS for master-eligible nodes. These settings ensure the cluster can handle data lake query loads while maintaining availability.

### IAM and Connector Permissions

Securing the pipeline requires careful IAM configuration. DynamoDB needs stream read permissions for change data capture, while Elasticsearch requires specific connector permissions. According to [`ElasticServiceAccounts.java`](https://github.com/elastic/elasticsearch/blob/main/ElasticServiceAccounts.java) (lines 49-58), service accounts must possess `cluster:admin/xpack/connector/*` permissions to execute Elastic Connectors. This permission set enables the connector to manage sync jobs and write to target indices securely.

## Integration Patterns: Connecting DynamoDB to Elasticsearch

For data lakes requiring both high-throughput storage and rich search capabilities, the recommended pattern stores raw events in DynamoDB while replicating searchable fields to Elasticsearch. The repository supports this through Elastic Connectors and custom pipeline implementations.

### Java Bulk Indexing Implementation

When building custom sync pipelines, use the Bulk API for efficient document ingestion:

```java
// DynamoDB client (v2)
DynamoDbClient ddb = DynamoDbClient.builder()
    .region(Region.US_EAST_1)
    .build();

// Scan the table (simple example – production should use pagination / parallel scans)
ScanResponse scan = ddb.scan(ScanRequest.builder()
    .tableName("my-lake-table")
    .build());

// Elasticsearch bulk request
BulkRequest.Builder bulk = new BulkRequest.Builder();

for (Map<String, AttributeValue> item : scan.items()) {
    // Convert DynamoDB map to JSON (using Jackson or manual mapping)
    String json = mapToJson(item);

    bulk.operations(op -> op
        .create(c -> c
            .index(i -> i
                .index("lake-index")
                .id(item.get("PK").s())
            )
            .document(json)
        )
    );
}

// Execute bulk request
ElasticsearchClient esClient = new ElasticsearchClient(
    new RestClientBuilder(HttpHost.create("https://my-es-domain.region.es.amazonaws.com"))
);
BulkResponse response = esClient.bulk(bulk.build());

// Handle errors
if (response.errors()) {
    response.items().stream()
        .filter(i -> i.error() != null)
        .forEach(i -> System.err.println("Failed: " + i.error().reason()));
}

```

The [`ElasticServiceAccounts.java`](https://github.com/elastic/elasticsearch/blob/main/ElasticServiceAccounts.java) file (lines 49-58) confirms that executing this bulk operation requires `cluster:admin/xpack/connector/*` permissions for connector-based service accounts.

### Elastic Connector Configuration

For managed synchronization, configure the AWS DynamoDB source connector. The sync-job lifecycle is documented in [`es-sync-rules.md`](https://github.com/elastic/elasticsearch/blob/main/es-sync-rules.md) (line 67):

```yaml

# file: dynamodb-connector.yml

service_type: aws-dynamodb
configuration:
  table: my-lake-table
  region: us-east-1
  access_key: ${AWS_ACCESS_KEY}
  secret_key: ${AWS_SECRET_KEY}
  scan_batch_size: 5000
  stream_enabled: true          # use DynamoDB Streams for incremental sync

  sync:
    schedule: "0 */5 * * * *"   # every 5 minutes

target_index:
  name: lake-index
  refresh: true
  mapping:
    properties:
      PK: { type: keyword }
      timestamp: { type: date }
      payload: { type: text }

```

### Terraform Deployment

Provision the Amazon Web Services Elasticsearch search engine for data lakes with optimized settings from [`cloud-aws-best-practices.md`](https://github.com/elastic/elasticsearch/blob/main/cloud-aws-best-practices.md) (lines 12-30):

```hcl
resource "aws_elasticsearch_domain" "lake_search" {
  domain_name = "lake-search"

  elasticsearch_version = "7.17"

  cluster_config {
    instance_type = "r5.large.elasticsearch"
    instance_count = 3
    zone_awareness_enabled = true
    zone_awareness_config {
      availability_zone_count = 2
    }
    // Use instance store for data nodes (recommended in cloud-aws-best-practices)
    storage_type = "instance"
  }

  ebs_options {
    ebs_enabled = false
  }

  zone_awareness_enabled = true

  advanced_options = {
    "rest.action.multi.allow_explicit_index" = "true"
  }

  // Enable fine-grained access control (required for connector IAM)
  advanced_security_options {
    enabled = true
    internal_user_database_enabled = true
  }
}

```

## Summary

- **DynamoDB** excels as a high-throughput, low-latency key-value store for raw data lake ingestion, offering predictable costs per request and automatic scaling.
- **Elasticsearch** (Amazon OpenSearch Service) provides the full-text search, aggregations, and semantic query capabilities required for analytics workloads, as implemented in the `elastic/elasticsearch` repository.
- **Hybrid architectures** leverage DynamoDB Streams or Elastic Connectors (documented in [`es-sync-rules.md`](https://github.com/elastic/elasticsearch/blob/main/es-sync-rules.md)) to maintain searchable indexes without duplicating storage costs.
- **AWS-specific optimizations** include using instance storage for data nodes and multi-AZ deployment, detailed in [`cloud-aws-best-practices.md`](https://github.com/elastic/elasticsearch/blob/main/cloud-aws-best-practices.md) (lines 12-30).
- **Security integration** requires granting `cluster:admin/xpack/connector/*` permissions to service accounts, as defined in [`ElasticServiceAccounts.java`](https://github.com/elastic/elasticsearch/blob/main/ElasticServiceAccounts.java) (lines 49-58).

## Frequently Asked Questions

### Can I use DynamoDB as the primary storage and Elasticsearch only for search indexes?

Yes, this is the recommended pattern for cost-efficient data lakes. Store immutable raw events in DynamoDB for cheap, scalable persistence, then replicate only the fields needed for search and analytics to Elasticsearch using DynamoDB Streams with Lambda or the Elastic Connector for DynamoDB. The connector's sync-job configuration is documented in [`es-sync-rules.md`](https://github.com/elastic/elasticsearch/blob/main/es-sync-rules.md) (line 67), allowing you to schedule incremental syncs every few minutes to keep indexes fresh without overloading your cluster.

### What permissions does Elasticsearch require to sync data from DynamoDB?

Elasticsearch service accounts need the `cluster:admin/xpack/connector/*` permission set to manage connector jobs and write to target indices. According to [`ElasticServiceAccounts.java`](https://github.com/elastic/elasticsearch/blob/main/ElasticServiceAccounts.java) (lines 49-58), these permissions enable the connector to execute sync rules, handle bulk indexing, and manage pipeline configurations. Additionally, the AWS IAM role running the connector requires `dynamodb:Scan`, `dynamodb:DescribeTable`, and `dynamodb:GetRecords` permissions for table access and stream consumption.

### How do I optimize Elasticsearch deployment costs on AWS for data lake workloads?

Deploy Elasticsearch with instance storage rather than EBS for data nodes to reduce storage costs while maintaining high throughput, as recommended in [`cloud-aws-best-practices.md`](https://github.com/elastic/elasticsearch/blob/main/cloud-aws-best-practices.md) (lines 12-30). Use multi-AZ deployments only for master-eligible nodes and critical data, and implement index lifecycle management to move older data to cheaper storage tiers. For data lakes specifically, maintain a smaller, curated search index in Elasticsearch while keeping raw historical data in DynamoDB or S3, minimizing the node count and instance sizes needed for your Elasticsearch cluster.

### What are the latency differences between querying DynamoDB versus Elasticsearch?

DynamoDB provides single-digit millisecond latency for point reads and writes when accessing items by primary key, making it ideal for real-time ingestion and lookup patterns. Elasticsearch delivers millisecond-level search latency for complex queries, but this depends on shard sizing, query complexity, and cluster load. For data lake architectures, use DynamoDB when you need immediate retrieval of specific records by ID, and leverage Elasticsearch when you require ad-hoc search, aggregations, or fuzzy matching across large datasets, as supported by the Query DSL and k-NN capabilities in the `elastic/elasticsearch` repository.