How Macro's Connection Gateway Manages WebSocket Connections at Scale

Macro's connection gateway scales WebSocket connections horizontally using AWS Fargate tasks behind an Application Load Balancer, leveraging target-tracking auto-scaling based on ALB request count per target to maintain approximately 1,000 concurrent connections per task while storing active sockets in an in-memory hashmap for real-time message routing.

Macro's connection gateway is a dedicated Rust micro-service within the macro-inc/macro repository designed to accept, manage, and route persistent WebSocket streams for real-time features like notifications, chat, and document syncing. The architecture decouples connection state from application logic, allowing the gateway to scale elastically based on demand without disrupting active client sessions.

Architecture Overview

The gateway runs as a stateless binary on AWS Fargate, positioned behind an Application Load Balancer (ALB) that terminates TLS and handles the HTTP to WebSocket protocol upgrade. When a client sends a request with an Upgrade: websocket header, the ALB forwards the raw TCP-level stream to a healthy Fargate task running the Rust connection-gateway binary.

This design places all connection-handling logic in the Rust service while relying on AWS infrastructure for load balancing and SSL termination. The separation ensures that scaling decisions are driven by infrastructure metrics rather than application state, enabling automatic horizontal scaling during traffic spikes.

Horizontal Auto-Scaling Strategy

The scaling implementation is defined in infra/stacks/connection-gateway/connection_gateway.ts and relies on multiple complementary mechanisms to handle varying load levels.

ALB Configuration and Idle Timeouts

The ALB is configured with an idleTimeout of 3600 seconds to prevent premature closure of idle WebSocket connections. This setting ensures that persistent sockets remain open even during periods of low message activity, which is critical for real-time applications where clients maintain long-lived connections.

The ALB supports unlimited concurrent connections subject to AWS service limits, forwarding each new WebSocket handshake to the healthiest available task. When load increases, the ALB distributes incoming handshakes across all running tasks based on the round-robin algorithm.

Fargate Auto-Scaling Policies

The setupAutoScaling method in connection_gateway.ts configures three distinct scaling policies:

  • Request-count scaling: Targets the ALBRequestCountPerTarget metric with a default value of 1000 connections per task, ensuring new Fargate instances spin up as handshake volume increases.
  • CPU-based scaling: Monitors task CPU utilization to catch computationally intensive message processing scenarios.
  • Memory-based scaling: Prevents out-of-memory conditions when the in-memory socket map grows large.

The request-count policy uses a target value of 1000, meaning AWS automatically adjusts the number of tasks to keep the average request count per target near this threshold. Scale-out cooldowns are set to 120 seconds, while scale-in cooldowns are 60 seconds to prevent thrashing.

Resource Allocation per Task

Each Fargate task is allocated 4 vCPU and 8 GiB of memory, providing sufficient resources to manage thousands of concurrent WebSocket connections in memory. This generous allocation allows a single task to handle the target 1,000 connections while leaving headroom for message processing spikes.

Tasks run the gateway binary alongside side-car containers for the Datadog agent and log router, ensuring observability without consuming the primary task's connection-handling capacity.

Connection Lifecycle and Routing

Understanding how connections flow through the system reveals why the stateless-task model works for stateful sockets.

WebSocket Upgrade Handling

When the ALB forwards a connection to a task, the Rust binary uses axum's WebSocketUpgrade handler to accept the stream, as implemented in services/connection_gateway/src/api/connection/mod.rs. After the upgrade completes, the socket is stored in an in-memory hashmap keyed by user ID, enabling direct message delivery without database lookups.

Because the socket map exists only in process memory, the gateway remains "stateless" from an infrastructure perspective—losing a task means losing its connections, but the ALB prevents new handshakes from reaching draining tasks while existing sockets finish naturally.

Inter-Service Message Routing

Application services communicate with the gateway through the WebSocketGatewayAdapter defined in crates/notification/src/outbound/websocket.rs. This adapter uses the ConnectionGatewayClient (crates/connection/src/outbound/connection_gateway_client.rs) to translate internal method calls into HTTP POST requests.

When code calls send_message(user_id, payload), the client POSTs to the gateway's internal API, which looks up the user ID in its hashmap and pushes the payload onto the appropriate WebSocket stream. This indirection allows backend services to remain agnostic about which specific Fargate task holds the user's connection.

Health Checks and Graceful Draining

The serviceLoadBalancer helper creates a target group with health checks configured against the /health endpoint. Tasks must pass this check to receive traffic; failures trigger automatic replacement.

When scaling down or during deployments, the ALB stops routing new handshakes to tasks marked for deregistration while allowing existing WebSocket connections to complete their sessions. This graceful draining ensures zero dropped messages during scaling events or rolling updates.

Observability and Alerting

The setupServiceAlarms function in connection_gateway.ts installs CloudWatch alarms for high CPU, memory utilization, and HTTP 5xx errors. These alarms trigger SNS notifications sent to the topic defined by CLOUD_TRAIL_SNS_TOPIC_ARN, alerting operators before the gateway becomes a bottleneck.

Datadog logging and metrics collection via side-car containers provide granular visibility into connection counts, message throughput, and latency per task.

Implementation Examples

Frontend Client Connection

Clients connect directly to the ALB endpoint using the Socket.IO protocol:

import { io } from 'socket.io-client';

const socket = io('wss://connection-gateway.prod.macro.inc', {
  transports: ['websocket'],
  reconnectionAttempts: 5,
});

socket.on('connect', () => console.log('WebSocket connected'));
socket.on('notification', (msg) => console.log('Notification:', msg));
socket.on('disconnect', () => console.warn('WebSocket closed'));

Backend Service Integration

Services send messages through the adapter layer without managing socket state:

use notification::outbound::websocket::{ConnectionGatewayClient, WebSocketGatewayAdapter};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let client = ConnectionGatewayClient::new("https://connection-gateway.prod.macro.inc")?;
    let realtime_sender = WebSocketGatewayAdapter::new(client);

    realtime_sender
        .send("user-123", serde_json::json!({ "type": "ping" }).to_string())
        .await?;

    Ok(())
}

Infrastructure as Code

The auto-scaling policy that drives horizontal scaling is defined as follows:

new aws.appautoscaling.Policy(
  `${BASE_NAME}-scaling-policy-request-count-${stack}`,
  {
    policyType: "TargetTrackingScaling",
    resourceId: serviceScalableTarget.resourceId,
    scalableDimension: serviceScalableTarget.scalableDimension,
    serviceNamespace: serviceScalableTarget.serviceNamespace,
    targetTrackingScalingPolicyConfiguration: {
      targetValue: 1000,
      predefinedMetricSpecification: {
        predefinedMetricType: "ALBRequestCountPerTarget",
        resourceLabel,
      },
      scaleInCooldown: 60,
      scaleOutCooldown: 120,
    },
  },
  { parent: this }
);

Summary

  • Macro's connection gateway uses AWS Fargate with Application Load Balancer to handle WebSocket upgrades and TLS termination.
  • Target-tracking auto-scaling maintains approximately 1,000 concurrent connections per task using the ALBRequestCountPerTarget metric.
  • The 4 vCPU/8 GiB task allocation supports thousands of simultaneous sockets stored in an in-memory hashmap keyed by user ID.
  • Graceful draining ensures existing connections complete naturally when tasks scale down or redeploy.
  • Backend services interact via the WebSocketGatewayAdapter and ConnectionGatewayClient to route messages without managing connection state.

Frequently Asked Questions

How does the gateway handle auto-scaling during sudden traffic spikes?

The gateway combines request-count scaling with CPU-based scaling policies to respond to load increases. The ALBRequestCountPerTarget metric triggers new task launches within 120 seconds when connection counts exceed 1,000 per task, while CPU alarms catch computationally heavy message processing scenarios that might not correlate directly with connection count.

What happens to existing WebSocket connections when a Fargate task terminates?

When a task enters a draining state, the ALB stops sending new WebSocket handshakes to that instance while maintaining existing TCP connections. The task continues processing messages for its current sockets until clients disconnect or the graceful timeout expires, ensuring zero interruption to active sessions during scaling events.

How do backend services route messages to specific users?

Services use the WebSocketGatewayAdapter (defined in crates/notification/src/outbound/websocket.rs), which wraps the ConnectionGatewayClient to send HTTP POST requests to the gateway's internal API. The gateway looks up the user ID in its in-memory hashmap and pushes the payload directly onto the corresponding WebSocket stream without requiring the backend to know which task hosts the connection.

Which source files control the scaling and connection behavior?

The infrastructure scaling logic resides in infra/stacks/connection-gateway/connection_gateway.ts, specifically the setupAutoScaling and setupServiceAlarms methods. Connection handling happens in services/connection_gateway/src/api/connection/mod.rs using axum, while message routing between services and the gateway is implemented in crates/connection/src/outbound/connection_gateway_client.rs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →