How Macro's Connection Gateway Handles WebSocket Connections at Scale

Macro's connection-gateway uses AWS Fargate with Application Load Balancer-backed auto-scaling to manage thousands of concurrent WebSocket connections through target-tracking policies that maintain roughly 1,000 connections per task.

The macro-inc/macro repository contains a dedicated connection-gateway microservice built to solve the hard problems of persistent WebSocket management: horizontal scaling, graceful draining, and reliable message routing. This deep dive examines how the gateway achieves production-grade scale without connection loss or service degradation.

Architecture Overview

The gateway sits at the edge of Macro's real-time infrastructure, handling notifications, chat, and document syncing. It runs as a Rust binary inside AWS Fargate containers, fronted by an Application Load Balancer (ALB) that terminates TLS and manages the HTTP-to-WebSocket upgrade dance.

The ALB configuration is critical to scale. In infra/stacks/connection-gateway/connection_gateway.ts, the idleTimeout is explicitly set to 3600 seconds—preventing premature closure of idle but healthy sockets that would otherwise force unnecessary reconnect storms.

Horizontal Scaling with Target-Tracking Auto-Scaling

Macro's scaling strategy centers on the setupAutoScaling method in connection_gateway.ts, which registers three complementary policies:

  • ALBRequestCountPerTarget — targets ~1,000 concurrent WebSocket handshakes per task
  • CPU utilization — catches compute-bound scenarios not reflected in connection count
  • Memory utilization — guards against the in-memory socket map consuming available RAM

When the request-count metric exceeds the target, Fargate launches new tasks within 120 seconds (scaleOutCooldown). During scale-in, the 60-second cooldown (scaleInCooldown) prevents thrashing while the ALB drains existing connections from tasks marked for termination.

new aws.appautoscaling.Policy(
  `${BASE_NAME}-scaling-policy-request-count-${stack}`,
  {
    policyType: "TargetTrackingScaling",
    resourceId: serviceScalableTarget.resourceId,
    scalableDimension: serviceScalableTarget.scalableDimension,
    serviceNamespace: serviceScalableTarget.serviceNamespace,
    targetTrackingScalingPolicyConfiguration: {
      targetValue: 1000,               // Aim for ~1 k hand‑shakes per task
      predefinedMetricSpecification: {
        predefinedMetricType: "ALBRequestCountPerTarget",
        resourceLabel,
      },
      scaleInCooldown: 60,
      scaleOutCooldown: 120,
    },
  },
  { parent: this }
);

Task Design for Connection Density

Each Fargate task receives 4 vCPU and 8 GiB memory—generous allocation that lets a single task absorb thousands of concurrent connections before triggering scale-out. Tasks run three containers:

  1. Primary: The Rust connection-gateway binary
  2. Sidecar: Datadog agent for metrics
  3. Sidecar: Log router for centralized logging

This resource headroom matters because the gateway maintains an in-memory hashmap of active sockets keyed by user ID. Unlike some distributed WebSocket solutions that externalize state to Redis, Macro keeps socket references local to maximize throughput and minimize latency for message fan-out.

WebSocket Upgrade and Connection Lifecycle

The upgrade flow begins in services/connection_gateway/src/api/connection/mod.rs, where axum's WebSocketUpgrade extractor handles the protocol switch:

  1. Client initiates wss://connection-gateway.<env>.macro.inc
  2. ALB detects Upgrade: websocket header, terminates TLS, forwards raw TCP
  3. Rust binary accepts via WebSocketUpgrade, stores socket in user-keyed map
  4. Application services push messages through ConnectionGatewayClient

The explicit statelessness of application logic—only the socket map holds ephemeral state—means tasks can be terminated without data loss. The ALB's graceful deregistration stops new handshakes to draining tasks while existing sockets complete naturally.

Message Routing Between Services

Backend services never touch WebSocket protocols directly. Instead, they use the WebSocketGatewayAdapter found in crates/notification/src/outbound/websocket.rs:

use notification::outbound::websocket::{ConnectionGatewayClient, WebSocketGatewayAdapter};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Build a client that knows the gateway’s HTTP endpoint
    let client = ConnectionGatewayClient::new("https://connection-gateway.prod.macro.inc")?;

    // Wrap it in the generic realtime sender used by the notification crate
    let realtime_sender = WebSocketGatewayAdapter::new(client);

    // Send a payload to user “user‑123”
    realtime_sender
        .send("user-123", serde_json::json!({ "type": "ping" }).to_string())
        .await?;

    Ok(())
}

The ConnectionGatewayClient in crates/connection/src/outbound/connection_gateway_client.rs translates these calls into internal HTTP POST requests that the gateway routes to the correct socket. This indirection decouples notification logic from transport concerns and enables testing with mocked adapters.

Health Monitoring and Alerting

Operational confidence comes from setupServiceAlarms in connection_gateway.ts, which provisions CloudWatch alarms for:

  • High CPU utilization — signals insufficient task count
  • High memory utilization — warns of socket map pressure
  • HTTP 5xx errors — catches gateway crashes or capacity exhaustion

Alarms publish to an SNS topic configured via CLOUD_TRAIL_SNS_TOPIC_ARN, ensuring human operators intervene before automated scaling proves insufficient.

Client Integration

Frontend code connects through standard Socket.IO, as shown in the repository's TypeScript examples:

import { io } from 'socket.io-client';

// The domain is generated by the Terraform stack:
//   connection-gateway.<env>.macro.inc
const socket = io('wss://connection-gateway.prod.macro.inc', {
  transports: ['websocket'],
  reconnectionAttempts: 5,
});

socket.on('connect', () => console.log('WebSocket connected'));
socket.on('notification', (msg) => console.log('Notification:', msg));
socket.on('disconnect', () => console.warn('WebSocket closed'));

The reconnectionAttempts configuration provides resilience during gateway scale-out events, when brief connection drops may occur as the ALB redistributes load.

Key Implementation Files

Component Source Path
Fargate service and auto-scaling definition infra/stacks/connection-gateway/connection_gateway.ts
Axum WebSocket upgrade handler services/connection_gateway/src/api/connection/mod.rs
Message handling logic services/connection_gateway/src/api/connection/messages.rs
Backend adapter for WebSocket sends crates/notification/src/outbound/websocket.rs
HTTP client for gateway communication crates/connection/src/outbound/connection_gateway_client.rs

Summary

  • Connection-gateway is a stateless-in-logic, stateful-in-memory Rust service running on AWS Fargate
  • ALB request-count auto-scaling targets 1,000 connections per task with 4 vCPU/8 GiB allocation
  • Graceful draining through ALB deregistration prevents connection loss during scale-in
  • Internal HTTP API decouples backend services from WebSocket complexity
  • Comprehensive CloudWatch alarms provide early warning before capacity limits

Frequently Asked Questions

How does Macro prevent WebSocket disconnections during deployments?

The ALB target group deregistration delay and Fargate's graceful shutdown hooks work together. When a task receives a termination signal, the ALB immediately stops routing new handshakes while existing connections drain naturally. The 3600-second idle timeout ensures quiet sockets aren't prematurely closed.

Why use in-memory socket storage instead of Redis?

Local hashmaps eliminate network roundtrips for message delivery. Since each socket is tied to a specific task and the ALB handles sticky routing through the handshake phase, external state would add latency without offering meaningful benefits for Macro's use case—messages are point-to-point, not broadcast.

What happens if a single user opens multiple WebSocket connections?

The user-keyed socket map in services/connection_gateway/src/api/connection/mod.rs supports multiple entries per key. All active sockets for a user receive the same message, with deduplication handled upstream in application logic if needed.

Can the gateway handle sudden traffic spikes beyond 1,000 connections per task?

Yes—the scaleOutCooldown of 120 seconds is conservative, but the ALB itself queues handshakes briefly while Fargate provisions new tasks. For predictable spikes (product launches, scheduled events), Macro likely pre-scales through scheduled auto-scaling actions, though this isn't visible in the analyzed source files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →