How to Achieve Horizontal Scaling with Load Balancers and Shared Storage for the Copilot SDK

To serve many concurrent users, deploy multiple Copilot CLI server instances behind a load balancer with a shared storage location for session state, allowing any instance to resume any session for even load distribution and high availability.

The github/copilot-sdk repository provides enterprise-grade infrastructure to scale Copilot services beyond a single instance. When traffic demands exceed the capacity of one server, implementing horizontal scaling with load balancers and shared storage for Copilot SDK ensures session continuity and fault tolerance across a distributed pool of CLI processes.

Architecture Overview

A horizontally scaled deployment consists of independent CLI servers behind a load balancer, with all instances pointing to a shared network filesystem. According to docs/setup/scaling.md, this topology allows any server in the pool to handle any user request because session data is decoupled from the compute nodes.

The standard configuration includes:

  • CLI Server Pool: Multiple processes listening on distinct ports (e.g., :4321, :4322, :4323)
  • Load Balancer: Distributes incoming traffic via round-robin or sticky-session algorithms
  • Shared Storage: Network File System (NFS) or cloud-mounted volumes at ~/.copilot/session-state/

The go/session_fs_provider.go file implements the filesystem-based session storage that makes this architecture possible. By placing this directory on shared storage, every CLI instance gains read/write access to the same session files.

Configure Shared Storage for Session State

Session persistence requires mounting a shared volume to the CLI's state directory. The Copilot CLI writes session data to ~/.copilot/session-state/, which must be accessible to every server in the pool.

For production deployments, mount persistent volumes using:

  • NFS: Traditional network-attached storage for on-premise clusters
  • Cloud File Systems: AWS EFS, Azure Files, or Google Cloud Filestore
  • Kubernetes PVCs: Persistent Volume Claims that mount to each pod

Ensure the filesystem supports concurrent read/write operations, as multiple CLI servers may access the same session files simultaneously during high-load scenarios.

Deploy CLI Servers in Headless Mode

Each CLI instance must run as a headless server binding to 0.0.0.0 so the load balancer can route traffic. Start servers with distinct ports as documented in docs/setup/backend-services.md:

copilot --headless --host 0.0.0.0 --port 4321

Backend SDK clients connect using RuntimeConnection.forUri():

import { CopilotClient, RuntimeConnection } from "@github/copilot-sdk";

const client = new CopilotClient({
    connection: RuntimeConnection.forUri("cli-server-1:4321"),
    mode: "empty",
});

Load Balancer Configuration Strategies

The scaling guide in docs/setup/scaling.md outlines two routing approaches:

Sticky Sessions (Session Affinity)

  • Pins a user to a specific CLI server using IP hashing or cookie-based routing
  • Eliminates the need for shared storage but risks uneven load distribution if users have variable activity levels

Shared Storage with Round-Robin

  • Routes requests to any available server via round-robin or least-connections algorithms
  • Requires networked filesystems but achieves true load balancing and high availability

For external reverse proxies, configure NGINX or HAProxy to forward localhost:<port> requests to the CLI pool. In Kubernetes, a Service resource automatically load-balances traffic across pods.

Client-Side Routing Logic

Your application can implement custom load balancing using the CLILoadBalancer pattern from the scaling guide. This TypeScript class provides both round-robin and sticky routing based on user ID hashing:

class CLILoadBalancer {
    private servers: string[];
    private idx = 0;

    constructor(servers: string[]) {
        this.servers = servers;
    }

    // Simple round-robin
    next(): string {
        const srv = this.servers[this.idx];
        this.idx = (this.idx + 1) % this.servers.length;
        return srv;
    }

    // Sticky routing based on userId
    forUser(userId: string): string {
        const hash = this.hashCode(userId);
        return this.servers[hash % this.servers.length];
    }

    private hashCode(s: string): number {
        let h = 0;
        for (let i = 0; i < s.length; i++) {
            h = (h << 5) - h + s.charCodeAt(i);
            h |= 0;
        }
        return Math.abs(h);
    }
}

Integrate the load balancer with your HTTP framework:

const lb = new CLILoadBalancer([
    "cli-1:4321",
    "cli-2:4321",
    "cli-3:4321",
]);

app.post("/chat", async (req, res) => {
    const server = lb.forUser(req.user.id);
    const client = new CopilotClient({
        connection: RuntimeConnection.forUri(server),
        mode: "empty",
    });
    
    const session = await client.createSession({
        sessionId: `user-${req.user.id}-${Date.now()}`,
        model: "gpt-5.4",
    });
    
    const reply = await session.sendAndWait({ prompt: req.body.message });
    res.json({ content: reply?.data.content });
});

Production Deployment Examples

Docker Compose with Shared Volume

version: "3.8"
services:
  copilot-cli:
    image: copilot-cli:latest
    command: ["--headless", "--host", "0.0.0.0", "--port", "4321"]
    environment:
      - COPILOT_GITHUB_TOKEN=${COPILOT_GITHUB_TOKEN}
    ports:
      - "4321:4321"
    volumes:
      - session-data:/root/.copilot/session-state

  api:
    build: .
    environment:
      - CLI_URL=copilot-cli:4321
    depends_on:
      - copilot-cli
    ports:
      - "3000:3000"

volumes:
  session-data:

Kubernetes Deployment with Persistent Storage

apiVersion: apps/v1
kind: Deployment
metadata:
  name: copilot-cli
spec:
  replicas: 3
  selector:
    matchLabels:
      app: copilot-cli
  template:
    metadata:
      labels:
        app: copilot-cli
    spec:
      containers:
      - name: copilot-cli
        image: your-registry/copilot-cli:latest
        args: ["--headless", "--host", "0.0.0.0", "--port", "4321"]
        env:
        - name: COPILOT_GITHUB_TOKEN
          valueFrom:
            secretKeyRef:
              name: copilot-secrets
              key: github-token
        ports:
        - containerPort: 4321
        volumeMounts:
        - name: session-state
          mountPath: /root/.copilot/session-state
      volumes:
      - name: session-state
        persistentVolumeClaim:
          claimName: copilot-sessions-pvc
---
apiVersion: v1
kind: Service
metadata:
  name: copilot-cli
spec:
  selector:
    app: copilot-cli
  ports:
  - port: 4321
    targetPort: 4321

Summary

  • Horizontal scaling requires running multiple CLI server instances behind a load balancer to distribute traffic across your Copilot SDK deployment.
  • Shared storage at ~/.copilot/session-state/ is mandatory for stateless routing, allowing any server to resume any session.
  • Headless mode enables the CLI to accept network connections via --headless --host 0.0.0.0 flags.
  • Client routing can be delegated to external load balancers or handled programmatically using the CLILoadBalancer pattern.
  • Production readiness demands persistent volumes, health checks, graceful shutdowns, and secret management for GitHub tokens.

Frequently Asked Questions

What is the difference between sticky sessions and shared storage for Copilot SDK?

Sticky sessions pin a specific user to a single CLI server instance for the duration of their session, which eliminates the need for shared filesystems but can lead to uneven load distribution. Shared storage allows any CLI server to handle any request, achieving true horizontal scaling but requiring a networked filesystem like NFS or cloud storage mounted at ~/.copilot/session-state/.

How should I secure GitHub tokens when scaling horizontally?

Store sensitive credentials like COPILOT_GITHUB_TOKEN in secret managers rather than environment variables or configuration files. In Kubernetes, use secretKeyRef to inject tokens from Secrets into pod environments. The docs/setup/multi-tenancy.md file provides additional guidance on per-user token isolation in multi-tenant deployments.

Can I use cloud storage instead of NFS for the session state directory?

Yes, cloud-native filesystems such as AWS EFS, Azure Files, or Google Cloud Filestore work as replacements for traditional NFS. Ensure the storage solution supports concurrent access and maintains POSIX file semantics, as the go/session_fs_provider.go implementation relies on standard filesystem operations.

What monitoring is required for a production horizontal scaling deployment?

Monitor active session counts per CLI instance, request latency from the load balancer, error rates on the /health endpoints, and disk I/O on the shared storage volume. The production checklist in docs/setup/scaling.md recommends implementing graceful shutdown handlers to allow in-flight requests to complete before terminating instances during scale-down events.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →