How MongoDB Docker Compose Entry Point Scripts Handle Startup Timing and Retries

The entry point scripts in the minhhungit/mongodb-cluster-docker-compose repository use a combination of initial sleep periods, local health-check loops, and bounded retry mechanisms to ensure MongoDB containers start in the correct order while tolerating transient network failures.

Orchestrating a distributed MongoDB sharded cluster with Docker Compose presents unique timing challenges. Containers must start sequentially—config servers first, then shards, finally the router—yet network latency and process initialization times vary. The custom Bash entry point scripts located in the scripts/ directory solve this by implementing explicit synchronization primitives rather than relying on Docker's depends_on alone.

Startup Sequence Overview

Each MongoDB component follows a three-phase startup protocol defined in its respective entry point script:

  1. Process Launch and Grace Period: The script starts mongod or mongos in the background, then pauses for sleep 5 to allow the process to initialize.
  2. Local Health Verification: A wait_for_mongo function polls the local instance using mongosh --eval "db.adminCommand('ping')" until it responds, guarded by kill -0 $MONGO_PID to detect process crashes.
  3. Peer Discovery with Bounded Retries: The script probes peer containers (e.g., configsvr02, shard01-b) using a retry loop with max_attempts=30 and a 2-second sleep between attempts.

Config Server Initialization

The scripts/entrypoint-configserver.sh script manages the config server replica set (rs-config-server). It must ensure all three config server nodes are reachable before attempting replica set initiation.

Local Instance Health Checks

After launching the MongoDB process, the script verifies local availability:


# Start mongod in background

mongod --configsvr --replSet rs-config-server ... &
MONGO_PID=$!

# Grace period for process initialization

sleep 5

# Poll until local instance responds or process dies

until mongosh --host 127.0.0.1 --eval "db.adminCommand('ping')" >/dev/null 2>&1; do
  if ! kill -0 $MONGO_PID 2>/dev/null; then
    echo "MongoDB process died unexpectedly"
    exit 1
  fi
  sleep 2
done

The kill -0 $MONGO_PID check prevents infinite loops if the mongod process crashes during startup.

Peer Discovery with Bounded Retries

Before initializing the replica set, the script verifies connectivity to the other config servers:


# Check peer availability with timeout

max_attempts=30
attempt=1

until mongosh --host configsvr02:27017 --eval "db.adminCommand('ping')" >/dev/null 2>&1; do
  if [ $attempt -eq $max_attempts ]; then
    echo "Timeout waiting for configsvr02"
    break
  fi
  echo "Waiting for configsvr02... (attempt $attempt/$max_attempts)"
  sleep 2
  attempt=$((attempt + 1))
done

This pattern repeats for configsvr03. If a peer never becomes ready within 30 attempts (approximately 60 seconds), the script logs a timeout but proceeds to attempt replica set initialization anyway, allowing the cluster to self-heal once all nodes are online.

Shard Replica Set Startup

The shard entry points (scripts/entrypoint-shard01.sh, scripts/entrypoint-shard02.sh, scripts/entrypoint-shard03.sh) follow an identical pattern to the config servers but target their respective replica sets (e.g., rs-shard-01).

Process Monitoring with kill -0

Each shard script monitors its local mongod process to prevent hanging during startup failures:

mongod --shardsvr --replSet rs-shard-01 ... &
MONGO_PID=$!
sleep 5

# Health check with crash detection

while ! mongosh --host 127.0.0.1 --eval "db.adminCommand('ping')" >/dev/null 2>&1; do
  if ! kill -0 $MONGO_PID 2>/dev/null; then
    echo "ERROR: MongoDB process crashed"
    exit 1
  fi
  sleep 2
done

Cross-Container Peer Polling

After confirming local readiness, the script polls the other two members of the shard replica set (e.g., shard01-b and shard01-c) using the same max_attempts=30 logic. This ensures the replica set can be initialized with a majority of nodes visible, though the script proceeds even if some peers are temporarily unreachable.

Router Service Coordination

The scripts/entrypoint-route.sh script manages the mongos router process, which requires all replica sets to have elected a primary before it can route queries correctly.

Primary Election Detection

Unlike the other scripts that check for process availability, the router script specifically waits for primary elections using rs.status():


# Wait for config server primary

until mongosh --host rs-config-server/configsvr01:27017,configsvr02:27017,configsvr03:27017 \
      --eval 'rs.status().members.some(m => m.stateStr === "PRIMARY")' | grep -q 'true'; do
  echo "Waiting for config server primary election..."
  sleep 5
done

This pattern repeats for each shard replica set (rs-shard-01, rs-shard-02, rs-shard-03). The script blocks until a primary is detected in each replica set, ensuring metadata consistency before exposing the router.

Mongos Launch Sequence

After verifying all primaries exist, the script starts mongos in the background and performs a final health check:


# Start mongos

mongos --configdb rs-config-server/configsvr01:27017,configsvr02:27017,configsvr03:27017 ... &
MONGOS_PID=$!

# Verify router responsiveness

until mongosh --port 27017 --eval 'db.adminCommand({ping:1})' >/dev/null 2>&1; do
  sleep 2
done

echo "Mongos router is ready"
wait $MONGOS_PID

The final wait $MONGOS_PID ensures the container remains running as long as the router process is active.

Summary

The MongoDB Docker Compose entry point scripts implement a robust, multi-layered approach to container startup synchronization:

  • Grace periods (sleep 5) allow processes to initialize before health checks begin.
  • Local health loops using mongosh ping commands verify instance readiness while kill -0 guards against process crashes.
  • Bounded retry mechanisms (max_attempts=30 with 2-second intervals) prevent infinite waits when peer containers are temporarily unreachable.
  • Primary election detection in the router script ensures metadata consistency before exposing the mongos service.
  • Process supervision via wait $MONGO_PID keeps containers alive after initialization completes.

These patterns ensure the sharded cluster starts reliably despite Docker Compose's lack of native startup ordering guarantees.

Frequently Asked Questions

How do the scripts prevent infinite loops if a MongoDB process crashes during startup?

Each script stores the MongoDB process ID in MONGO_PID immediately after launching it in the background. Inside the local health-check loops, the scripts execute kill -0 $MONGO_PID to verify the process still exists. If this check fails—indicating the process has terminated—the script prints an error message and exits with a non-zero status rather than continuing to poll indefinitely.

Why does the router script wait for primary elections instead of just checking if containers are running?

The mongos router requires access to a primary node in each replica set to read and write sharding metadata. Simply verifying that container processes are alive does not guarantee a replica set has completed election and chosen a primary—especially during initial startup when nodes are still negotiating. The entrypoint-route.sh script specifically queries rs.status() and checks for stateStr === "PRIMARY" to ensure the cluster is fully operational before exposing the router service.

What happens if a peer container never becomes ready within the 30-attempt limit?

If a peer container fails to respond to mongosh ping commands after 30 attempts (approximately 60 seconds), the script logs a timeout message such as "Timeout waiting for configsvr02" and breaks out of the retry loop. Crucially, the script does not exit; it proceeds to attempt replica set initialization anyway. This design allows the cluster to self-heal once the delayed node eventually comes online, as MongoDB's native replication protocols will synchronize it with the existing primary.

How does the initial sleep period contribute to startup reliability?

The sleep 5 command inserted immediately after launching mongod provides a brief grace period for the MongoDB process to initialize its internal data structures, bind to the network interface, and begin accepting connections. Without this pause, the subsequent mongosh health checks might execute before the server is ready, causing unnecessary retry cycles or false crash detections. This sleep acts as a coarse synchronization barrier that compensates for process startup latency in containerized environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →