Part 4 · 1 chapters · ~8 min

Elasticsearch and OpenSearch Clusters

Indexes, mappings and dynamic mapping pitfalls, primary and replica shards and routing, the query and fetch phases, deep pagination and search_after, aliases and zero-downtime reindexing, keeping the index in sync with the database (CDC), cluster health and sizing, and the Elasticsearch and OpenSearch split.

5

Clusters, shards and aliases

code
PUT merchants_v8 { ...new settings and mappings... }
POST _reindex { "source": { "index": "merchants_v7" }, "dest": { "index": "merchants_v8" } }
POST _aliases { "actions": [ { "remove": { "index": "merchants_v7", "alias": "merchants" } },
                             { "add":    { "index": "merchants_v8", "alias": "merchants" } } ] }   // atomic switch
GET _cluster/health ; GET _cat/shards?v ; GET _cat/indices?v

// deep pages: search_after with a sort tiebreaker instead of from=10000
GET merchants/_search { "size": 50, "sort": [ { "_score": "desc" }, { "id": "asc" } ], "search_after": [3.21, "m_8812"] }

Keeping in sync: the database stays the source of truth; changes flow to the index through CDC or an outbox (Scaling course P7), and the index can always be rebuilt with a reindex. Dynamic mapping guesses field types from the first document it sees; define explicit mappings for production indexes. Elasticsearch and OpenSearch (the AWS-led fork from 2021) share the core concepts.

AN ELASTICSEARCH / OPENSEARCH CLUSTER
an index split into shards, each a Lucene index, replicated across nodes
clientcoordinating nodescatter, gather, mergeshard 0 (primary)node Ashard 1 (primary)node Bshard 2 (primary)node Creplica of 0node Balias merchants → merchants_v7cluster managermetadata, allocation
swipe the figure sideways, or tap expand for full screen
1/5
shards
An index is split into primary shards, each a complete Lucene index. Documents are routed by hash of their id (or a routing key). The number of primary shards is fixed at creation.
primary shards = Lucene indexesfixed at creation