How to Implement Nested Aggregation Queries for Hierarchical Elasticsearch Facets

Nested aggregation queries in Elasticsearch isolate child documents using the path parameter, creating execution scopes where sub-aggregations build hierarchical facets by grouping on fields within nested objects.

Nested aggregation queries enable faceted navigation across complex hierarchical data structures in Elasticsearch. According to the elastic/elasticsearch source code, the platform stores nested objects as separate hidden documents linked to their parent _source, requiring the nested aggregation type to establish proper execution contexts for sub-aggregations on these child documents.

Understanding Nested Document Storage

According to server/src/main/java/org/elasticsearch/search/NestedUtils.java, Elasticsearch stores each nested object as an independent hidden document while maintaining the parent-child relationship. This internal representation allows the engine to query nested fields independently, but standard aggregations cannot cross the parent-child boundary without explicit nesting.

Defining Nested Field Mappings

Before executing nested aggregation queries, you must declare fields as nested type in your index mapping. This configuration instructs Elasticsearch to index nested objects as separate documents while preserving their relationship to the parent document.

PUT /products
{
  "mappings": {
    "properties": {
      "resellers": {
        "type": "nested",
        "properties": {
          "reseller": { "type": "keyword" },
          "price":    { "type": "double" }
        }
      }
    }
  }
}

The type: nested declaration is mandatory; omitting it causes the "Nested aggregation requires a path" error because the aggregation expects a nested field path.

Building Hierarchical Facet Structures

A hierarchical facet consists of nested aggregation queries that create tree-like drill-down capabilities. The pattern uses a nested aggregation as the root bucket to isolate child documents, followed by terms or range sub-aggregations that group on fields within those nested objects.

GET /products/_search?size=0
{
  "aggs": {
    "by_reseller": {
      "nested": { "path": "resellers" },
      "aggs": {
        "reseller_name": {
          "terms": { "field": "resellers.reseller" },
          "aggs": {
            "price_ranges": {
              "range": {
                "field": "resellers.price",
                "ranges": [
                  { "to": 200 },
                  { "from": 200, "to": 500 },
                  { "from": 500 }
                ]
              }
            }
          }
        }
      }
    }
  }
}

The outer nested aggregation defined in NestedAggregationBuilder.java establishes the execution scope, while inner aggregations compute facets specific to the nested context.

Multi-Level Aggregation Trees

For deeper hierarchies, chain multiple sub-aggregation levels. The following example demonstrates a three-level facet: author → genre → price histogram.

GET /library/_search?size=0
{
  "aggs": {
    "nested_books": {
      "nested": { "path": "books" },
      "aggs": {
        "by_author": {
          "terms": { "field": "books.author", "size": 10 },
          "aggs": {
            "by_genre": {
              "terms": { "field": "books.genre", "size": 5 },
              "aggs": {
                "price_histogram": {
                  "histogram": {
                    "field": "books.price",
                    "interval": 50
                  }
                }
              }
            }
          }
        }
      }
    }
  }
}

This structure returns buckets where each author contains genre buckets, which in turn contain price distribution histograms.

Internal Architecture and Execution

The implementation of nested aggregation queries relies on specialized components that manage the transition from parent to child document contexts.

When a nested aggregation executes, NestedAggregatorFactory (defined in server/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregatorFactory.java) instantiates NestedAggregator to handle the runtime logic. The NestedAggregator class (located at server/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregator.java) opens a child SearchContext that points specifically to the nested document IDs, ensuring sub-aggregations run only against those hidden child documents.

Key Implementation Components

During execution, the nested aggregator rewrites global ordinals for the nested field, allowing standard aggregations like terms, range, and histogram to function exactly as they do on flat fields while remaining scoped to child documents.

Performance Optimization Strategies

Optimizing nested aggregation queries requires managing memory usage and bucket cardinality across distributed shards.

  • shard_size parameter: Increase this value on inner terms aggregations to ensure accurate bucket counts across shards, preventing missing buckets in high-cardinality nested fields.
  • collect_mode: The default breadth_first collection mode caches top-level documents for child aggregations, reducing memory pressure with high-cardinality fields. For shallow hierarchies with low cardinality, depth_first may improve performance.
  • composite aggregation: For extremely high-cardinality nested fields, replace terms aggregations with composite aggregations to paginate through results safely without exhausting memory.

Troubleshooting Common Pitfalls

Symptom Cause Solution
Missing buckets in results terms.size or default shard_size too low for shard distribution Increase size and explicitly set shard_size to a higher value
"Nested aggregation requires a path" error Field not mapped as nested type Redefine the field with "type": "nested" in the index mapping
Excessive memory consumption breadth_first mode caching large document sets Reduce size parameters or switch to depth_first for shallow hierarchies

Summary

  • Nested aggregation queries require fields mapped with "type": "nested" to establish proper execution contexts.
  • The path parameter in NestedAggregationBuilder isolates child documents for sub-aggregation processing.
  • Internal components including NestedAggregator and NestedUtils manage parent-child document ID translation and context isolation.
  • Hierarchical facets are constructed by chaining sub-aggregations under a root nested aggregation.
  • Performance depends on configuring shard_size and collect_mode appropriate to your data cardinality.

Frequently Asked Questions

Why do I need to use nested aggregations instead of standard terms aggregations?

Standard aggregations operate at the root document level and cannot access fields within nested objects independently. Without a nested aggregation, Elasticsearch treats the entire array of nested objects as a single entity, causing incorrect facet counts when multiple nested values exist in one document. The nested aggregation type creates an isolated execution context using the path parameter, ensuring sub-aggregations calculate statistics against individual nested documents rather than parent documents.

How does the path parameter work in nested aggregation queries?

The path parameter specifies which nested field to enter for the aggregation scope. As implemented in NestedAggregationBuilder.java, this parameter validates against the mapping to ensure the field is properly typed as nested, then instructs NestedAggregator to open a child SearchContext restricted to documents matching that path. This mechanism effectively switches the aggregation context from parent documents to their associated hidden child documents stored by NestedUtils.

Can I use reverse_nested aggregations to climb back up the hierarchy?

Yes. After drilling down into nested documents, the reverse_nested aggregation allows you to return to the parent document context for additional aggregations. This is useful when you need to aggregate on both nested fields and parent-level fields within the same query tree. The aggregation appears as a sibling to other sub-aggregations under the nested bucket, creating bidirectional navigation capabilities in your hierarchical facets.

What causes memory issues with nested aggregations and how do I fix them?

Memory pressure typically results from the breadth_first collection mode (default for high-cardinality fields) caching large sets of top-level documents while processing child aggregations. According to the Elasticsearch source, this mode builds bitsets of matching parent documents to optimize child aggregation execution. To resolve this, reduce the size parameter on your aggregations to limit bucket generation, or explicitly set collect_mode: depth_first if your hierarchy depth is shallow and cardinality is manageable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →