What Is an Index in Elasticsearch? Architecture and Source Code Deep Dive

An index in Elasticsearch is a logical namespace that stores a collection of documents with shared mappings and settings, physically implemented by an Index object (containing a name and UUID) and an IndexMetadata object that tracks shard configuration, versioning, and lifecycle state in the cluster state.

In the elastic/elasticsearch repository, an index represents the fundamental unit of data organization, but it is not merely a container—it is a distributed, versioned metadata construct managed by the cluster state. Understanding how indices are defined at the source code level is essential for optimizing cluster operations, troubleshooting allocation issues, and designing scalable data architectures.

Core Components of an Elasticsearch Index

The Index Class: Logical Identity

According to the source code in server/src/main/java/org/elasticsearch/index/Index.java, the Index class is a lightweight, immutable value object that encapsulates two critical identifiers: the human-readable index name and a globally unique UUID generated at creation time. This dual-identifier system ensures that indices can be tracked consistently across logs, routing tables, and persistence layers even if the index name is later aliased or modified. The class's toString() method returns the format [name/uuid], which appears throughout Elasticsearch logs and debugging output.

IndexMetadata: The Persistent Configuration

While the Index class provides identity, IndexMetadata—defined in server/src/main/java/org/elasticsearch/cluster/metadata/IndexMetadata.java—contains the complete, immutable snapshot of an index's configuration. This includes the number of primary shards, replica count, mappings, settings, aliases, routing filters, and lifecycle policy associations. Because it implements Diffable<IndexMetadata>, the class supports efficient cluster state updates by computing diffs rather than serializing the entire object on every change.

Architectural Characteristics of an Index

Immutable Metadata and Versioning

Every modification to an index—whether updating settings, changing mappings, or adjusting allocation filters—generates a new IndexMetadata instance through methods like withIncrementedVersion(), withInSyncAllocationIds(), or withSettings(). The object maintains multiple version fields including version, mappingVersion, settingsVersion, aliasesVersion, and transportVersion, which enable nodes to reconcile state during rolling upgrades and verify compatibility via indexCompatibilityVersion().

Shard Routing and Allocation

The IndexMetadata drives shard placement decisions through fields such as routingNumShards, routingFactor, routingPaths, and allocation filters (require, include, exclude). These properties are consulted by allocation deciders including DiskThresholdDecider and ShardsLimitAllocationDecider to determine which nodes host specific shards. The routing configuration is immutable once defined, ensuring consistent hash-based routing of documents to shards.

Lifecycle and System Flags

Indices may carry special flags stored in IndexMetadata: isSystem marks system indices reserved for internal Elasticsearch operations, while isHidden prevents indices from appearing in wildcard queries by default. Additionally, Index Lifecycle Management (ILM) data including lifecyclePolicyName and lifecycleExecutionState resides here, enabling automated rollover, shrink, and deletion operations.

Practical Implementation: Creating and Managing Indices

Create an Index via REST API

The following cURL command creates an index named my-index with custom shard configuration and mappings:

curl -X PUT "localhost:9200/my-index" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "number_of_shards": 3,
    "number_of_replicas": 1
  },
  "mappings": {
    "properties": {
      "timestamp": { "type": "date" },
      "message":   { "type": "text" }
    }
  }
}'

This request triggers the creation of an Index object (with a generated UUID) and an IndexMetadata entry stored in the cluster state.

Create an Index Using the Java High-Level REST Client

For Java applications, the High-Level REST Client provides a type-safe builder pattern that mirrors the REST API structure:

RestHighLevelClient client = new RestHighLevelClient(
    RestClient.builder(new HttpHost("localhost", 9200, "http")));

CreateIndexRequest request = new CreateIndexRequest("my-index");
request.settings(
    Settings.builder()
        .put("index.number_of_shards", 3)
        .put("index.number_of_replicas", 1)
);
request.mapping(
    "{ \"properties\": { \"timestamp\": { \"type\": \"date\" }, \"message\": { \"type\": \"text\" } } }",
    XContentType.JSON
);

CreateIndexResponse response = client.indices().create(request, RequestOptions.DEFAULT);
System.out.println("Created index UUID: " + response.index());

Behind the scenes, the client sends a JSON payload identical to the cURL example; the master node builds an IndexMetadata instance using the builder pattern seen throughout the source.

Retrieving Index Metadata

To inspect the configuration stored in IndexMetadata, use the Get Index API:

GetIndexRequest getRequest = new GetIndexRequest("my-index");
GetIndexResponse getResponse = client.indices().get(getRequest, RequestOptions.DEFAULT);
System.out.println("Settings: " + getResponse.getSettings().getAsMap());
System.out.println("Mappings: " + getResponse.getMappings());

The returned data mirrors what is stored in IndexMetadata, including settings, mappings, aliases, and other flags.

Key Source Files

Understanding the index abstraction requires familiarity with these core files in the elastic/elasticsearch repository:

Summary

  • An index in Elasticsearch is a logical namespace comprising an Index object (name and UUID) and IndexMetadata (configuration and schema).
  • The Index class in Index.java provides immutable identity, while IndexMetadata in IndexMetadata.java stores shard counts, mappings, settings, and lifecycle data.
  • All metadata is immutable; updates create new instances with incremented version fields to support safe concurrent modifications and rolling upgrades.
  • Shard routing and allocation are driven by fields within IndexMetadata, consulted by allocation deciders to determine node placement.
  • Indices support system and hidden flags plus Index Lifecycle Management (ILM) policies stored directly in the metadata.

Frequently Asked Questions

What is the difference between an index and a shard in Elasticsearch?

An index is the logical namespace that groups documents with shared mappings and settings, represented by the Index and IndexMetadata classes. A shard is the physical Lucene index that actually stores the data; each index is split into one or more primary shards (plus replicas) distributed across the cluster. The IndexMetadata defines how many shards exist and how they are routed, but the shards themselves are managed by the IndexShard class in server/src/main/java/org/elasticsearch/index/shard/IndexShard.java.

How does Elasticsearch ensure index metadata consistency across the cluster?

Elasticsearch stores IndexMetadata in the cluster state, which is maintained by the master node and replicated to all data nodes. Because IndexMetadata implements Diffable<IndexMetadata>, the system computes diffs (deltas) rather than broadcasting the entire state on every change. When you update an index setting or mapping, the master creates a new IndexMetadata instance with an incremented version field, computes the diff, and publishes it to the cluster. Nodes apply these diffs atomically to their local cluster state, ensuring consistency.

Can an index name be changed after creation in Elasticsearch?

You cannot rename an index directly because the Index object's name is immutable and serves as the key in the cluster state's metadata map. However, you can create an alias that points to the existing index, then remove the old name if needed. The IndexMetadata stores these aliases in its aliases field (a map of AliasMetadata objects). When you search using an alias, Elasticsearch resolves it to the underlying index UUID (stored in the Index class), ensuring that data is retrieved from the correct physical shards regardless of which alias name you use.

What are system indices in Elasticsearch and how are they marked?

System indices are internal indices used by Elasticsearch features like security, monitoring, and machine learning to store configuration or state data. They are marked by the isSystem boolean field inside IndexMetadata. When this flag is true, the index is protected from user modifications and wildcard queries unless specifically allowed. Similarly, the isHidden flag prevents indices from appearing in wildcard searches by default. Both flags are stored in the immutable IndexMetadata and are set during index creation via the index.hidden setting and system index descriptors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →