Elasticsearch vs Prometheus for Container and Server Monitoring: Key Differences and Performance Metrics
Elasticsearch excels at long-term storage and complex analytics across logs and metrics using its document-based data streams, while Prometheus specializes in high-frequency scraping, PromQL calculations, and cloud-native alerting, with the Elasticsearch repository providing a Prometheus Remote-Write integration to bridge both worlds.
When evaluating monitoring solutions for containerized environments and server infrastructure, operators often compare the Elastic Stack's metric capabilities against Prometheus's time-series architecture. The elastic/elasticsearch repository contains a dedicated Prometheus Remote-Write integration (x-pack/plugin/prometheus) that enables Prometheus to push metrics directly into Elasticsearch data streams, demonstrating how these systems can complement each other. Understanding their architectural differences and performance characteristics is essential for selecting the right tool—or hybrid architecture—for your observability needs.
Architectural Differences
The fundamental divergence between Elasticsearch and Prometheus begins with data ingestion models, storage formats, and query capabilities.
| Aspect | Elasticsearch Monitoring Stack | Prometheus |
|---|---|---|
| Data Ingestion | Agents like Metricbeat or the Prometheus Remote-Write plugin (PrometheusRemoteWriteRestAction) receive metrics via HTTP/HTTPS and index them into metrics data streams following the naming convention metrics-<dataset>.prometheus-<namespace>. |
Pull-based scrapes over HTTP or remote-write pushes of protobuf-encoded samples to a receiver endpoint. |
| Storage Format | Documents stored in Elasticsearch data streams; each sample is a JSON/CBOR document with @timestamp, labels, and a field named after the metric value, as constructed in PrometheusRemoteWriteTransportAction.buildIndexRequest. |
Time-series stored in a column-oriented TSDB; each series identified by labels, with values stored in compressed blocks. |
| Query Language | Elasticsearch Query DSL supporting full-text search, aggregations, and pipeline processors. Enables ad-hoc analytics, joins, and machine learning on metrics. | PromQL – a functional query language specialized for time-series calculations; limited to time-window operations. |
| Alerting | Integrated with Watcher and the Alerting framework; alerts based on any Elasticsearch query. | Built-in Alertmanager; alerts derived from PromQL expressions. |
| Scalability | Sharded indices across the cluster; horizontal scaling by adding nodes, each holding a subset of shards managed by PrometheusIndexTemplateRegistry. |
Horizontal scaling via federation or remote-write to a long-term store; each Prometheus server handles its own scrape targets. |
| Retention | Index Lifecycle Policies (ILM) control rollover, shrink, and deletion; supports raw samples for days and roll-ups for months. | Time-based retention per series; down-sampling optional via remote-write or external tools. |
| Security | TLS, authentication, and RBAC baked into Elasticsearch; settings governed by xpack.monitoring.* configuration files as documented in monitoring-settings.md. |
TLS and basic auth on remote-write endpoints; RBAC delegated to the receiving system (e.g., the Elasticsearch plugin). |
Critical Performance Metrics to Evaluate
When benchmarking these systems for container and server monitoring, focus on these operational characteristics:
| Metric | Elasticsearch | Prometheus |
|---|---|---|
| Ingestion Latency | Measured from HTTP request receipt to write acknowledgement (typically milliseconds). The plugin logs the number of time-series and samples per request via logger.debug in PrometheusRemoteWriteTransportAction.doExecute. |
Typically sub-millisecond for local scrapes; remote-write adds network latency and protobuf parsing overhead. |
| Throughput (samples/sec) | Depends on bulk size; the plugin builds a BulkRequest and executes it in one shot via bulkRequestBuilder.execute. Throughput scales with write threads and shard distribution. |
Limited by scrape interval and number of targets; remote-write can handle millions of samples/sec if the receiver (Elasticsearch) can keep up. |
| CPU Utilization | CPU spent parsing CBOR, indexing, and running aggregations. Indexing overhead rises with field cardinality (e.g., many unique label combinations). | CPU mainly used for scraping, metric exposition, and remote-write protobuf encoding. |
| Memory Usage | Index buffers, field data, and the JVM heap. High cardinality labels increase memory pressure. | In-memory TSDB buffers (chunks) plus a fixed-size ring buffer for recent samples. |
| Storage Efficiency | Depends on mapping: numeric fields store efficiently, but each distinct label becomes a separate keyword field, potentially increasing index size. | Uses highly compressed delta-encoding; efficient for high-cardinality numeric series, but each label set is stored as a separate series. |
| Query Latency | Aggregation latency grows with number of shards and cardinality of label fields; mitigated by roll-ups or pre-aggregated indices. | PromQL query latency depends on time range and number of series; typically fast for recent data, slower for long ranges. |
| Operational Overhead | Requires a JVM-based cluster, ILM policies, and monitoring of node health. | Runs as a single binary; needs separate Alertmanager and possibly remote-write storage for long-term retention. |
Selecting the Right Tool for Your Environment
Use Elasticsearch when:
- You need rich ad-hoc analytics across logs and metrics in a single platform.
- You already run the Elastic Stack (Kibana, Beats) and want unified RBAC and UI.
- Long-term retention, roll-ups, and integration with machine learning are required.
Use Prometheus when:
- You prefer a pull-based model with simple service discovery.
- Low-latency alerting based on time-series calculations is paramount.
- You want a lightweight, container-native collector without a JVM.
Hybrid Approach:
Deploy Prometheus for short-term scrapes and forward the data via remote-write to Elasticsearch. The plugin exposes the /_prometheus/api/v1/write endpoint, allowing Prometheus to push metrics into Elasticsearch data streams. This gives you Prometheus-style collection with Elasticsearch's long-term analytics capabilities.
Implementing the Prometheus Remote-Write Integration
Configuring Prometheus Remote-Write to Elasticsearch
To send metrics from Prometheus to Elasticsearch, configure the remote_write section in prometheus.yml:
global:
scrape_interval: 15s
remote_write:
- url: "https://elastic-node:9200/_prometheus/api/v1/write"
basic_auth:
username: "elastic"
password: "changeme"
# TLS configuration can be added under tls_config
The endpoint is defined in PrometheusRemoteWriteRestAction and accepts protobuf payloads (application/x-protobuf). The plugin validates the dataset and namespace parameters, then creates an index request per sample via PrometheusRemoteWriteRestAction.prepareRequest.
Direct Protobuf Submission
For testing or custom clients, you can post a protobuf payload directly:
# Generate binary protobuf from RemoteWrite.WriteRequest
# (see proto file under x-pack/plugin/prometheus/proto/RemoteWrite.proto)
cat <<EOF > write.pb
# binary protobuf content
EOF
curl -X POST \
-H "Content-Type: application/x-protobuf" \
--data-binary @write.pb \
"https://elastic-node:9200/_prometheus/myapp/api/v1/write?dataset=myapp&namespace=prod"
The request routes to PrometheusRemoteWriteRestAction, which builds a BulkRequest using BulkRequestBuilder and stores each sample in a data stream named metrics-myapp.prometheus-prod as implemented in PrometheusRemoteWriteTransportAction.buildIndexRequest.
Querying Metrics in Elasticsearch
Once stored, metrics become fields in Elasticsearch documents. Query them using the Elasticsearch Query DSL:
GET /metrics-myapp.prometheus-prod-*/_search
{
"size": 0,
"aggs": {
"avg_cpu": {
"avg": { "field": "cpu_usage_seconds_total" }
}
},
"query": {
"range": { "@timestamp": { "gte": "now-5m" } }
}
}
Because each metric becomes a field (e.g., cpu_usage_seconds_total), you can use any Elasticsearch aggregation to compute averages, percentiles, or histograms.
Enabling the Prometheus Plugin
Activate the index template registry that creates data streams automatically by adding to elasticsearch.yml:
xpack.prometheus.enabled: true
xpack.prometheus.registry.enabled: true
These settings instantiate the PrometheusPlugin class and enable PROMETHEUS_REGISTRY_ENABLED functionality, allowing the system to automatically manage the metrics-*.prometheus data stream templates via PrometheusIndexTemplateRegistry.
Summary
- Elasticsearch stores Prometheus metrics as documents in data streams (
metrics-*.prometheus-*), enabling complex aggregations and long-term retention via ILM policies, but requires JVM resources and careful management of high-cardinality labels. - Prometheus excels at high-frequency scraping, sub-millisecond local storage, and PromQL calculations, but relies on remote-write integrations like the Elasticsearch plugin (
PrometheusRemoteWriteRestAction) for long-term analytics. - The Prometheus Remote-Write integration in
x-pack/plugin/prometheusallows hybrid architectures where Prometheus handles collection and Elasticsearch handles storage, with metrics indexed viaBulkRequestBuilderand queryable through the standard Elasticsearch Query DSL. - Key performance differentiators include ingestion latency (milliseconds vs sub-millisecond), storage efficiency (document-based vs column-oriented TSDB), and query capabilities (full-text aggregations vs time-series functions).
Frequently Asked Questions
What is the Prometheus Remote-Write integration in Elasticsearch?
The Prometheus Remote-Write integration is an X-Pack plugin located in x-pack/plugin/prometheus that exposes an HTTP endpoint (/_prometheus/api/v1/write) allowing Prometheus servers to push time-series data directly into Elasticsearch. Implemented in PrometheusRemoteWriteRestAction and PrometheusRemoteWriteTransportAction, the plugin accepts protobuf-encoded samples, validates dataset and namespace parameters, and constructs BulkRequest objects to index documents into metrics-*.prometheus-* data streams.
How does storage efficiency compare between Elasticsearch and Prometheus for metrics?
Prometheus uses a column-oriented TSDB with delta-encoding compression, making it highly efficient for high-cardinality numeric time-series where each unique label set constitutes a separate series. Elasticsearch stores each sample as a JSON/CBOR document in a data stream, where each metric becomes a field and each label becomes a keyword field; this offers flexibility for complex aggregations but increases storage overhead and mapping complexity when handling high-cardinality container labels.
Which query language should I use for container monitoring metrics?
Use PromQL when you need specialized time-series calculations, rate functions, and instant vectors for alerting on recent container performance data within Prometheus. Use Elasticsearch Query DSL when you require ad-hoc analytics, full-text search across logs and metrics, joins between datasets, or machine learning on long-term historical data stored in Elasticsearch data streams. The Elasticsearch plugin allows you to query Prometheus-derived metrics using either approach depending on where the data resides.
What are the key performance metrics to monitor when running Prometheus with Elasticsearch remote-write?
Monitor ingestion latency from Prometheus remote-write to Elasticsearch, which should remain in the low milliseconds as measured from HTTP request receipt to write acknowledgement in PrometheusRemoteWriteTransportAction.doExecute. Track throughput in samples per second, noting that Elasticsearch scales via BulkRequest batching while Prometheus handles millions of samples per second when properly configured. Finally, watch storage growth and JVM heap pressure in Elasticsearch, as high-cardinality container labels can rapidly expand the index mapping and memory usage compared to Prometheus's compressed TSDB chunks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →