How to Perform a Python Elasticsearch Search for Large Datasets Without Hitting Memory Limits

Use the elasticsearch.helpers.scan generator, which wraps the server-side scroll API to stream documents in batches rather than loading the entire result set into memory at once.

When executing a python elasticsearch search against indices containing millions of documents, standard request patterns quickly exhaust RAM by attempting to deserialize all hits simultaneously. The elastic/elasticsearch repository solves this through the scroll API, a server-side cursor mechanism that transmits data in configurable chunks. By leveraging the high-level helpers.scan utility, you can retrieve exhaustive result sets while maintaining constant memory usage regardless of total dataset size.

Why Standard Pagination Fails for Large Results

The default from/size pagination approach is unsuitable for bulk extraction. Elasticsearch enforces an index.max_result_window limit of 10,000 hits by default, and exceeding this threshold requires the server to materialize and sort all preceding documents for every request. This creates linear slowdowns and exponential memory pressure on both the cluster and your Python process. For exhaustive retrieval, you need a streaming architecture that avoids materializing the full result set.

How the Scroll API Streams Results

The scroll API maintains a lightweight search context on the Elasticsearch node, allowing the client to pull documents sequentially without re-executing the query. This mechanism is implemented in the Java server code through several coordinated components:

The Scroll Lifecycle

The workflow follows three distinct phases:

  1. Initial search – Returns the first batch of hits alongside a _scroll_id token.
  2. Subsequent scroll requests – Each request transmits the current _scroll_id and receives the next batch; the server may rewrite the ID for load-balancing purposes.
  3. Clear scroll – After the final batch, or upon error, the client issues a clear request to destroy the server-side context.

Because each batch is processed independently before requesting the next, only a small slice of the total dataset resides in Python memory at any moment.

Implementing Efficient Large Dataset Retrieval in Python

Using helpers.scan for Automatic Streaming

The helpers.scan function abstracts the scroll lifecycle into a Python generator, automatically handling the initial search, pagination, and final cleanup.

from elasticsearch import Elasticsearch, helpers

es = Elasticsearch("http://localhost:9200")

def stream_large_dataset(index_name, search_query):
    """
    Lazily yields every document matching the query without loading
    the full result set into RAM.
    """
    for doc in helpers.scan(
        client=es,
        index=index_name,
        query=search_query,
        scroll="5m",          # Keep context alive for 5 minutes between requests

        size=1000,            # Documents per batch (tune based on document size)

        _source=True,         # Return full source; use False + stored_fields for specific fields

    ):
        yield doc["_source"]

# Usage: process millions of documents with constant memory

for source in stream_large_dataset("logs-2024", {"match_all": {}}):
    process_document(source)

Key implementation details:

  • scroll="5m" sets the keep-alive window; if your processing logic between iterations exceeds this duration, the scroll expires.
  • size=1000 controls the batch granularity—larger values reduce network round-trips but increase per-batch memory consumption.
  • Automatic cleanup occurs when the generator is exhausted or garbage collected, invoking clear_scroll via RestClearScrollAction to prevent resource leaks.

Manual Scroll Control

For scenarios requiring dynamic keep-alive adjustment or mid-stream error handling, use the low-level scroll API directly.


# Initial search with scroll context

resp = es.search(
    index="logs-2024",
    body={"query": {"match_all": {}}},
    scroll="2m",
    size=5000
)

scroll_id = resp["_scroll_id"]

try:
    while resp["hits"]["hits"]:
        for hit in resp["hits"]["hits"]:
            process_document(hit["_source"])
        
        # Fetch next batch using the scroll ID

        resp = es.scroll(scroll_id=scroll_id, scroll="2m")
        scroll_id = resp.get("_scroll_id")
        
finally:
    # Critical: Clear the scroll context even if processing fails

    if scroll_id:
        es.clear_scroll(scroll_id=scroll_id)

This manual implementation mirrors the server-side flow: the initial request creates a SearchScrollRequest, subsequent calls route through RestSearchScrollAction with InternalScrollSearchRequest objects, and cleanup invokes RestClearScrollAction.

Alternative: Point-in-Time with Search-After

For use cases requiring a consistent snapshot without maintaining a long-lived scroll context, the point-in-time (PIT) API combined with search_after provides similar streaming capabilities. The Python client offers helpers.pit_scan for this pattern. However, for pure retrieval workloads without concurrent indexing, the standard scroll API remains the most memory-efficient method.

Summary

  • Scroll API maintains a server-side cursor to avoid loading millions of documents into Python memory simultaneously.
  • helpers.scan automates the scroll lifecycle, including the final clear_scroll call to prevent resource leaks.
  • Batch size (the size parameter) directly controls memory usage per iteration—tune this based on your average document size.
  • Java implementation files including SearchScrollRequest.java, InternalScrollSearchRequest.java, and RestClearScrollAction.java handle the transport and cleanup operations transparently.
  • Never use from/size for deep pagination beyond the max_result_window (default 10,000), as it forces the server to materialize and sort all preceding hits.

Frequently Asked Questions

What is the maximum number of documents I can retrieve using helpers.scan?

There is no fixed upper limit. The generator yields documents until the query result set is exhausted, regardless of whether the index contains thousands or billions of documents. Memory usage remains constant based on your configured size parameter rather than total result count.

The scroll parameter (e.g., "5m") specifies how long the server retains the search context between requests. If your Python processing time between consuming batches exceeds this window, Elasticsearch discards the context and raises an exception on the next scroll request. Set this value based on your expected per-batch processing latency.

Why is from/size pagination inefficient for large datasets?

The from/size method requires the server to materialize, sort, and skip all documents preceding your target window. This creates O(n) complexity that becomes exponentially slower as you paginate deeper, and it hard stops at the index.max_result_window setting (default 10,000). Scroll avoids this by maintaining a live cursor that references the query execution state directly.

Do I need to manually clear the scroll context when using helpers.scan?

No. The helpers.scan generator automatically invokes the clear scroll action when the iterator is exhausted or when an exception propagates out of the generator. However, if you manually break out of the loop early using break, you should explicitly call es.clear_scroll to release server resources immediately.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →