How to Implement Nested Aggregation Queries for Hierarchical Elasticsearch Facets
Nested aggregation queries in Elasticsearch isolate child documents using the path parameter, creating execution scopes where sub-aggregations build hierarchical facets by grouping on fields within nested objects.
Nested aggregation queries enable faceted navigation across complex hierarchical data structures in Elasticsearch. According to the elastic/elasticsearch source code, the platform stores nested objects as separate hidden documents linked to their parent _source, requiring the nested aggregation type to establish proper execution contexts for sub-aggregations on these child documents.
Understanding Nested Document Storage
According to server/src/main/java/org/elasticsearch/search/NestedUtils.java, Elasticsearch stores each nested object as an independent hidden document while maintaining the parent-child relationship. This internal representation allows the engine to query nested fields independently, but standard aggregations cannot cross the parent-child boundary without explicit nesting.
Defining Nested Field Mappings
Before executing nested aggregation queries, you must declare fields as nested type in your index mapping. This configuration instructs Elasticsearch to index nested objects as separate documents while preserving their relationship to the parent document.
PUT /products
{
"mappings": {
"properties": {
"resellers": {
"type": "nested",
"properties": {
"reseller": { "type": "keyword" },
"price": { "type": "double" }
}
}
}
}
}
The type: nested declaration is mandatory; omitting it causes the "Nested aggregation requires a path" error because the aggregation expects a nested field path.
Building Hierarchical Facet Structures
A hierarchical facet consists of nested aggregation queries that create tree-like drill-down capabilities. The pattern uses a nested aggregation as the root bucket to isolate child documents, followed by terms or range sub-aggregations that group on fields within those nested objects.
GET /products/_search?size=0
{
"aggs": {
"by_reseller": {
"nested": { "path": "resellers" },
"aggs": {
"reseller_name": {
"terms": { "field": "resellers.reseller" },
"aggs": {
"price_ranges": {
"range": {
"field": "resellers.price",
"ranges": [
{ "to": 200 },
{ "from": 200, "to": 500 },
{ "from": 500 }
]
}
}
}
}
}
}
}
}
The outer nested aggregation defined in NestedAggregationBuilder.java establishes the execution scope, while inner aggregations compute facets specific to the nested context.
Multi-Level Aggregation Trees
For deeper hierarchies, chain multiple sub-aggregation levels. The following example demonstrates a three-level facet: author → genre → price histogram.
GET /library/_search?size=0
{
"aggs": {
"nested_books": {
"nested": { "path": "books" },
"aggs": {
"by_author": {
"terms": { "field": "books.author", "size": 10 },
"aggs": {
"by_genre": {
"terms": { "field": "books.genre", "size": 5 },
"aggs": {
"price_histogram": {
"histogram": {
"field": "books.price",
"interval": 50
}
}
}
}
}
}
}
}
}
}
This structure returns buckets where each author contains genre buckets, which in turn contain price distribution histograms.
Internal Architecture and Execution
The implementation of nested aggregation queries relies on specialized components that manage the transition from parent to child document contexts.
When a nested aggregation executes, NestedAggregatorFactory (defined in server/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregatorFactory.java) instantiates NestedAggregator to handle the runtime logic. The NestedAggregator class (located at server/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregator.java) opens a child SearchContext that points specifically to the nested document IDs, ensuring sub-aggregations run only against those hidden child documents.
Key Implementation Components
NestedAggregationBuilder: Validates thepathparameter and constructs the aggregation factory. Located inserver/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregationBuilder.java.NestedAggregatorFactory: Creates the runtime aggregator instance and manages the sub-aggregation pipeline. Found inserver/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregatorFactory.java.NestedAggregator: Opens a childSearchContextusing document ID sets and executes sub-aggregations against nested documents only. Implemented inserver/src/main/java/org/elasticsearch/search/aggregations/bucket/nested/NestedAggregator.java.NestedUtils: Translates parent-child document IDs and buildsFixedBitSetstructures for efficient nested document identification. Defined inserver/src/main/java/org/elasticsearch/search/NestedUtils.java.NestedIdentity: Exposes the nested path information when retrieving individual hits, enabling result interpretation. Located inserver/src/main/java/org/elasticsearch/search/SearchHit.java.
During execution, the nested aggregator rewrites global ordinals for the nested field, allowing standard aggregations like terms, range, and histogram to function exactly as they do on flat fields while remaining scoped to child documents.
Performance Optimization Strategies
Optimizing nested aggregation queries requires managing memory usage and bucket cardinality across distributed shards.
shard_sizeparameter: Increase this value on innertermsaggregations to ensure accurate bucket counts across shards, preventing missing buckets in high-cardinality nested fields.collect_mode: The defaultbreadth_firstcollection mode caches top-level documents for child aggregations, reducing memory pressure with high-cardinality fields. For shallow hierarchies with low cardinality,depth_firstmay improve performance.compositeaggregation: For extremely high-cardinality nested fields, replacetermsaggregations withcompositeaggregations to paginate through results safely without exhausting memory.
Troubleshooting Common Pitfalls
| Symptom | Cause | Solution |
|---|---|---|
| Missing buckets in results | terms.size or default shard_size too low for shard distribution |
Increase size and explicitly set shard_size to a higher value |
"Nested aggregation requires a path" error |
Field not mapped as nested type |
Redefine the field with "type": "nested" in the index mapping |
| Excessive memory consumption | breadth_first mode caching large document sets |
Reduce size parameters or switch to depth_first for shallow hierarchies |
Summary
- Nested aggregation queries require fields mapped with
"type": "nested"to establish proper execution contexts. - The
pathparameter inNestedAggregationBuilderisolates child documents for sub-aggregation processing. - Internal components including
NestedAggregatorandNestedUtilsmanage parent-child document ID translation and context isolation. - Hierarchical facets are constructed by chaining sub-aggregations under a root
nestedaggregation. - Performance depends on configuring
shard_sizeandcollect_modeappropriate to your data cardinality.
Frequently Asked Questions
Why do I need to use nested aggregations instead of standard terms aggregations?
Standard aggregations operate at the root document level and cannot access fields within nested objects independently. Without a nested aggregation, Elasticsearch treats the entire array of nested objects as a single entity, causing incorrect facet counts when multiple nested values exist in one document. The nested aggregation type creates an isolated execution context using the path parameter, ensuring sub-aggregations calculate statistics against individual nested documents rather than parent documents.
How does the path parameter work in nested aggregation queries?
The path parameter specifies which nested field to enter for the aggregation scope. As implemented in NestedAggregationBuilder.java, this parameter validates against the mapping to ensure the field is properly typed as nested, then instructs NestedAggregator to open a child SearchContext restricted to documents matching that path. This mechanism effectively switches the aggregation context from parent documents to their associated hidden child documents stored by NestedUtils.
Can I use reverse_nested aggregations to climb back up the hierarchy?
Yes. After drilling down into nested documents, the reverse_nested aggregation allows you to return to the parent document context for additional aggregations. This is useful when you need to aggregate on both nested fields and parent-level fields within the same query tree. The aggregation appears as a sibling to other sub-aggregations under the nested bucket, creating bidirectional navigation capabilities in your hierarchical facets.
What causes memory issues with nested aggregations and how do I fix them?
Memory pressure typically results from the breadth_first collection mode (default for high-cardinality fields) caching large sets of top-level documents while processing child aggregations. According to the Elasticsearch source, this mode builds bitsets of matching parent documents to optimize child aggregation execution. To resolve this, reduce the size parameter on your aggregations to limit bucket generation, or explicitly set collect_mode: depth_first if your hierarchy depth is shallow and cardinality is manageable.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →