How to Scale Copilot CLI Horizontally with Load Balancers: A Complete Guide

GitHub Copilot CLI operates as a stateless JSON-RPC server over TCP, allowing you to run multiple instances behind a TCP load balancer by leveraging automatic port discovery and connection token validation.

GitHub Copilot CLI (part of the copilot-sdk repository) can be deployed as a horizontally scalable service thanks to its stateless architecture and TCP-based JSON-RPC protocol. Because the CLI maintains no persistent state on disk and communicates via standard TCP streams, you can distribute traffic across multiple CLI processes using any TCP-capable load balancer. This guide explains the architecture, connection handling, and implementation details needed to scale Copilot CLI horizontally in production environments.

Architecture Principles for Horizontal Scaling

The Copilot CLI is designed as a stateless JSON-RPC server that communicates over TCP (or optionally stdio). This architecture makes horizontal scaling straightforward because the server does not persist any user-specific state to disk.

Stateless Design

Each CLI process keeps all session state in memory only. According to the official guidance in docs/setup/scaling.md, the CLI does not write to disk, making it safe to terminate or restart any instance without data loss. This statelessness means any backend instance can serve any client request, eliminating the need for sticky sessions or shared storage between instances.

Shared Configuration Across Instances

If you need identical experiment assignments, skill directories, or telemetry settings on every instance, pass them via the SDK’s ClientOptions. Parameters such as experimentAssignments, skillDirectories, and otelConfig are serialized into the JSON shape the CLI expects, ensuring uniform behavior across your load-balanced pool. This configuration is applied at spawn time, so all CLI instances start with identical settings.

Deploying Multiple CLI Instances

To scale horizontally, you must spawn multiple CLI processes and register them with your load balancer. The SDKs automate port discovery, token generation, and connection management.

TCP Mode Activation

While the CLI supports stdio for single-instance use, TCP mode is required for horizontal scaling. Launch the CLI with --port <port> to bind to a TCP socket and accept raw JSON-RPC streams. Setting the port to 0 instructs the OS to dynamically allocate a free port, which is essential when spawning multiple instances on the same host.

Port Discovery Mechanism

When launched with port 0, the CLI prints PORT=<num> on stdout once the socket is bound. The SDKs monitor this output to capture the actual listening port before proceeding with registration.

In rust/src/lib.rs, the Client::spawn_tcp method handles this port-wait logic, parsing the stdout stream until it detects the port announcement. Similarly, the Node.js implementation in nodejs/src/client.ts contains the spawnTcp function that waits for the "TCP port wait complete" log line before returning control to your application.

Connection Token Security

When an SDK spawns its own CLI instance, it automatically generates a UUID token unless you provide one explicitly. This token must be included in the JSON-RPC connect request, and the CLI validates it before processing subsequent commands. This mechanism prevents unauthorized processes from hijacking TCP connections in a multi-tenant environment.

As implemented in the Rust client (Client::start in rust/src/lib.rs) and the Node.js client (token field in ClientOptions in nodejs/src/client.ts), token generation happens automatically when you set the connection parameter to None or null.

Graceful Shutdown Handling

When a CLI process exits, the SDK automatically closes the TCP socket and removes the backend from the balancer's pool. This ensures clients do not receive half-open connections during scaling events, rolling deployments, or process restarts.

Load Balancer Configuration Requirements

Proper load balancer configuration ensures optimal performance and reliability when scaling Copilot CLI.

TCP vs HTTP Routing

Configure your load balancer to operate in TCP mode (Layer 4) rather than HTTP mode (Layer 7). The Copilot CLI uses raw JSON-RPC streams over TCP, not HTTP requests. HTTP-aware load balancers may buffer, modify, or terminate the stream, breaking the JSON-RPC protocol. Use raw TCP forwarding to preserve message framing.

Health Check Strategy

Implement TCP health checks that attempt a simple JSON-RPC ping to verify backend health. Because the protocol is stateless, a failed health check indicates the specific instance should be removed from the pool immediately. Configure your orchestration system to spawn replacement instances dynamically using the SDK's port discovery mechanism when health checks fail.

SDK-Specific Implementation Patterns

Each official SDK provides utilities to spawn TCP-mode CLI instances and handle port discovery.

Rust Implementation

The Rust SDK in rust/src/lib.rs provides the Client::spawn_tcp method for launching CLI instances:

use copilot_sdk::Client;
use copilot_sdk::ClientOptions;

let opts = ClientOptions::default()
    .with_mode(copilot_sdk::mode::ClientMode::CopilotCli)
    .with_tcp_port(0)               // 0 → OS‑assigned port
    .with_connection_token(None);   // SDK will generate a UUID

let client = Client::new(opts).await?;
let port = client.tcp_port().expect("TCP mode must expose a port");

// Register `port` (and the autogenerated token) with your load balancer here

Node.js Implementation

The Node.js client in nodejs/src/client.ts exposes spawnTcp and runtime information:

import { CopilotClient, RuntimeConnection } from '@github/copilot-sdk';

const conn = RuntimeConnection.forTcp({ port: 0 }); // OS picks a free port
const client = new CopilotClient({ runtimeConnection: conn });

await client.start();                     // waits for port announcement
const { port, token } = client.runtimeInfo; // expose to load balancer

Python Implementation

Python developers use RuntimeConnection.for_tcp as shown in python/copilot/client.py and documented in python/README.md:

from copilot import client, runtime

# Create a TCP runtime; port=0 lets the OS pick a free port

runtime_conn = runtime.RuntimeConnection.for_tcp(port=0)

c = client.CopilotClient(runtime_connection=runtime_conn)
c.start()               # blocks until the CLI prints its port

port = c.runtime_info.port
token = c.runtime_info.connection_token

# Now add `port`/`token` to your load‑balancer pool

Java Implementation

The Java SDK in java/src/main/java/com/github/copilot/rpc/CopilotClientOptions.java provides getters and setters for TCP configuration:

CopilotClientOptions opts = new CopilotClientOptions()
    .setMode(ClientMode.COPILOT_CLI)
    .setTcpPort(0)                    // let OS allocate
    .setConnectionToken(null);        // SDK will auto‑generate

CopilotClient client = new CopilotClient(opts);
client.start();                      // waits for port announcement
int port = client.getTcpPort();
String token = client.getConnectionToken();

Summary

  • Stateless architecture: Copilot CLI maintains no disk state, enabling safe horizontal scaling across multiple instances behind a load balancer.
  • TCP mode requirement: Use TCP transport (not stdio or HTTP) when spawning CLI instances to ensure compatibility with standard load balancers.
  • Automatic port discovery: Set port to 0 and let the SDK parse the PORT=<num> stdout announcement from the CLI process.
  • Security via tokens: Auto-generated UUID tokens in the JSON-RPC connect request prevent unauthorized access to CLI instances.
  • Load balancer configuration: Use Layer 4 TCP load balancing with JSON-RPC ping health checks, removing failed instances automatically when connections close.
  • Consistent configuration: Pass experimentAssignments, skillDirectories, and otelConfig via ClientOptions to ensure uniform behavior across all instances.

Frequently Asked Questions

Does Copilot CLI maintain session state that affects load balancing?

No. Copilot CLI is completely stateless and does not write user data to disk. All session information resides in memory, meaning any instance can handle any request. This design eliminates the need for sticky sessions or shared session stores when configuring your load balancer.

What load balancer algorithms work best with Copilot CLI?

Round-robin or least-connections algorithms work well because the JSON-RPC protocol is stateless and uniform. Avoid algorithms that assume HTTP semantics or require session affinity. Ensure your load balancer supports raw TCP forwarding (Layer 4) rather than HTTP proxying (Layer 7) to prevent protocol corruption.

How do I handle CLI process failures in a load-balanced setup?

Monitor TCP connection health using JSON-RPC ping requests. When a CLI process exits, the SDK closes the socket automatically, causing health checks to fail. Configure your load balancer to remove failed backends immediately and spawn replacement instances dynamically using the SDK's port discovery mechanism.

Can I use HTTP/HTTPS load balancers instead of TCP?

No. Copilot CLI uses raw JSON-RPC over TCP, not HTTP. Using an HTTP-aware load balancer will corrupt the protocol stream or cause connection failures. Always configure Layer 4 TCP forwarding to preserve the binary JSON-RPC message framing between the client and CLI instances.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →