How to Perform an Elasticsearch Delete Document Operation by Field Value

The most efficient way to perform an elasticsearch delete document operation based on a specific field value is to use the Delete-by-Query API (_delete_by_query), which executes a distributed scroll-and-bulk deletion across shards without requiring individual document IDs.

When working with the elastic/elasticsearch repository, developers often need to remove documents matching specific criteria rather than known IDs. While the standard Delete API requires an exact _id, performing an elasticsearch delete document operation by field value demands a different approach that leverages the cluster's distributed query engine and bulk processing capabilities.

How the Delete-by-Query API Works in Elasticsearch

The Delete-by-Query implementation follows a multi-layered architecture that transforms a query-based request into efficient bulk deletions.

REST Layer Entry Point

The HTTP request POST /{index}/_delete_by_query is handled by RestDeleteByQueryAction located at modules/reindex/src/main/java/org/elasticsearch/reindex/RestDeleteByQueryAction.java. This class parses the JSON request body, constructs a DeleteByQueryRequest object, and forwards it to the transport layer for cluster-wide execution.

Transport Layer Execution

At modules/reindex/src/main/java/org/elasticsearch/reindex/TransportDeleteByQueryAction.java, the transport action receives the request and creates a BulkByScrollTask. It initiates an AsyncDeleteByQueryAction that runs asynchronously across all relevant shards. This component manages the lifecycle of the deletion task, including cancellation support and progress tracking.

Bulk Deletion Engine

The actual deletion logic resides in AsyncDeleteByQueryAction (modules/reindex/src/main/java/org/elasticsearch/reindex/AsyncDeleteByQueryAction.java). This class implements the scroll-and-bulk pattern:

  1. Executes a scroll query to fetch document IDs matching the field value in batches
  2. Constructs bulk delete requests from each batch
  3. Submits bulk operations to the underlying index engine
  4. Repeats until all matching documents are processed

This approach minimizes network round-trips and leverages Elasticsearch's internal optimizations, including soft deletes and per-shard parallel execution.

Why Delete-by-Query is More Efficient Than Single-ID Deletes

When comparing approaches for field-based deletion, the Delete-by-Query API provides significant performance advantages over iterating single-document deletes:

  • Discovery Efficiency: Uses the inverted index to locate matching documents on each shard without requiring upfront knowledge of _id values, unlike RestDeleteAction (server/src/main/java/org/elasticsearch/rest/action/document/RestDeleteAction.java) which requires an exact ID.
  • Batch Processing: Groups deletions into bulk requests per shard, reducing per-document overhead compared to individual DeleteRequest objects (server/src/main/java/org/elasticsearch/action/delete/DeleteRequest.java).
  • Parallel Execution: Each shard processes its slice of matching documents concurrently, maximizing cluster resource utilization.
  • Soft-Delete Awareness: Integrates with the engine's soft-delete mechanism for safe recovery and versioning, managed through the BulkByScrollTask infrastructure.

Implementation Examples

HTTP API (cURL)

Delete all documents where the user field equals "bob":

curl -X POST "http://localhost:9200/my-index/_delete_by_query?refresh=true" -H 'Content-Type: application/json' -d '
{
  "query": {
    "term": {
      "user": "bob"
    }
  }
}'

The POST endpoint maps to RestDeleteByQueryAction, while the term query restricts the operation to matching documents. The refresh=true parameter forces a near-real-time refresh, making deletions immediately visible.

Java High-Level REST Client

import org.elasticsearch.client.RequestOptions;
import org.elasticsearch.client.RestHighLevelClient;
import org.elasticsearch.index.query.TermQueryBuilder;
import org.elasticsearch.index.reindex.BulkByScrollResponse;
import org.elasticsearch.index.reindex.DeleteByQueryRequest;

// client is an already-configured RestHighLevelClient
DeleteByQueryRequest request = new DeleteByQueryRequest("my-index");
request.setQuery(new TermQueryBuilder("user", "bob"));
request.setRefresh(true);                 // optional: make changes visible instantly
request.setConflicts("proceed");          // handle version conflicts gracefully

BulkByScrollResponse response = client.deleteByQuery(request, RequestOptions.DEFAULT);
System.out.println("Deleted docs: " + response.getDeleted());

Behind the scenes, DeleteByQueryRequest is serialized and sent to RestDeleteByQueryAction, which forwards to TransportDeleteByQueryAction and ultimately AsyncDeleteByQueryAction for execution.

Single-ID Delete Alternative

When you already know the document ID, use the standard Delete API for marginally better performance on individual documents:

import org.elasticsearch.action.delete.DeleteRequest;
import org.elasticsearch.action.delete.DeleteResponse;
import org.elasticsearch.client.RequestOptions;
import org.elasticsearch.client.RestHighLevelClient;

DeleteRequest request = new DeleteRequest("my-index", "doc-id-123");
DeleteResponse response = client.delete(request, RequestOptions.DEFAULT);
System.out.println("Result: " + response.getResult());

This request is handled by RestDeleteAction (server/src/main/java/org/elasticsearch/rest/action/document/RestDeleteAction.java), which creates a DeleteRequest and invokes TransportDeleteAction directly, bypassing the query phase entirely.

Key Source Files in the Elasticsearch Codebase

Component File Path
REST entry point for Delete-by-Query RestDeleteByQueryAction.java modules/reindex/src/main/java/org/elasticsearch/reindex/RestDeleteByQueryAction.java
Transport implementation for Delete-by-Query TransportDeleteByQueryAction.java modules/reindex/src/main/java/org/elasticsearch/reindex/TransportDeleteByQueryAction.java
Request object for Delete-by-Query DeleteByQueryRequest.java server/src/main/java/org/elasticsearch/index/reindex/DeleteByQueryRequest.java
Bulk delete engine (async) AsyncDeleteByQueryAction.java modules/reindex/src/main/java/org/elasticsearch/reindex/AsyncDeleteByQueryAction.java
REST entry point for single-ID delete RestDeleteAction.java server/src/main/java/org/elasticsearch/rest/action/document/RestDeleteAction.java
Delete request object (single ID) DeleteRequest.java server/src/main/java/org/elasticsearch/action/delete/DeleteRequest.java

Summary

  • The Delete-by-Query API (_delete_by_query) is the most efficient method for performing an elasticsearch delete document operation based on field values, as it eliminates the need to fetch and iterate individual document IDs.
  • The implementation uses a scroll-and-bulk pattern via AsyncDeleteByQueryAction, processing documents in batches across shards to minimize network overhead.
  • Single-ID deletes handled by RestDeleteAction are only optimal when the document ID is already known, bypassing the query phase but requiring explicit identification of each document.
  • The architecture separates concerns between REST handling (RestDeleteByQueryAction), transport coordination (TransportDeleteByQueryAction), and asynchronous execution (AsyncDeleteByQueryAction), enabling scalable, parallel deletion across the cluster.

Frequently Asked Questions

What is the difference between Delete-by-Query and the Delete API?

The Delete API requires an exact document _id and is handled by RestDeleteAction, making it suitable for targeted single-document removal when you know the identifier. Delete-by-Query, handled by RestDeleteByQueryAction, accepts a query to locate matching documents across shards, making it ideal for bulk deletions based on field values without requiring individual IDs.

Does Delete-by-Query support script-based deletion?

Yes, the Delete-by-Query API supports script-based filtering through the script parameter in the request body. When provided, TransportDeleteByQueryAction passes the script to AsyncDeleteByQueryAction, which evaluates the script against each matching document during the scroll phase to determine whether it should be deleted, enabling complex conditional logic beyond simple field matching.

How does Elasticsearch handle version conflicts during Delete-by-Query?

The Delete-by-Query API handles version conflicts through the conflicts parameter, which defaults to aborting the operation when a conflict is detected. When set to "proceed" as shown in the Java example, AsyncDeleteByQueryAction continues processing subsequent documents and returns a BulkByScrollResponse containing a count of version conflicts encountered, allowing applications to audit which documents were modified concurrently.

Is Delete-by-Query atomic or can it be partially applied?

Delete-by-Query is not atomic; it can be partially applied across the cluster. The operation processes documents in batches using the scroll API, and each batch is deleted independently. If the operation fails partway through, some documents will be deleted while others remain, though the BulkByScrollResponse reports exactly how many documents were successfully deleted before the failure occurred.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →