# How to Scale Copilot CLI Horizontally with Load Balancers: A Complete Guide

> Scale Copilot CLI horizontally with load balancers. Discover how to run multiple instances behind a TCP load balancer using automatic port discovery and connection token validation.

- Repository: [GitHub/copilot-sdk](https://github.com/github/copilot-sdk)
- Tags: how-to-guide
- Published: 2026-08-02

---

**GitHub Copilot CLI operates as a stateless JSON-RPC server over TCP, allowing you to run multiple instances behind a TCP load balancer by leveraging automatic port discovery and connection token validation.**

GitHub Copilot CLI (part of the copilot-sdk repository) can be deployed as a horizontally scalable service thanks to its stateless architecture and TCP-based JSON-RPC protocol. Because the CLI maintains no persistent state on disk and communicates via standard TCP streams, you can distribute traffic across multiple CLI processes using any TCP-capable load balancer. This guide explains the architecture, connection handling, and implementation details needed to scale Copilot CLI horizontally in production environments.

## Architecture Principles for Horizontal Scaling

The Copilot CLI is designed as a **stateless** JSON-RPC server that communicates over **TCP** (or optionally stdio). This architecture makes horizontal scaling straightforward because the server does not persist any user-specific state to disk.

### Stateless Design

Each CLI process keeps all session state in memory only. According to the official guidance in [`docs/setup/scaling.md`](https://github.com/github/copilot-sdk/blob/main/docs/setup/scaling.md), the CLI does not write to disk, making it safe to terminate or restart any instance without data loss. This statelessness means any backend instance can serve any client request, eliminating the need for sticky sessions or shared storage between instances.

### Shared Configuration Across Instances

If you need identical experiment assignments, skill directories, or telemetry settings on every instance, pass them via the SDK’s `ClientOptions`. Parameters such as `experimentAssignments`, `skillDirectories`, and `otelConfig` are serialized into the JSON shape the CLI expects, ensuring uniform behavior across your load-balanced pool. This configuration is applied at spawn time, so all CLI instances start with identical settings.

## Deploying Multiple CLI Instances

To scale horizontally, you must spawn multiple CLI processes and register them with your load balancer. The SDKs automate port discovery, token generation, and connection management.

### TCP Mode Activation

While the CLI supports stdio for single-instance use, **TCP mode is required for horizontal scaling**. Launch the CLI with `--port <port>` to bind to a TCP socket and accept raw JSON-RPC streams. Setting the port to `0` instructs the OS to dynamically allocate a free port, which is essential when spawning multiple instances on the same host.

### Port Discovery Mechanism

When launched with port `0`, the CLI prints `PORT=<num>` on stdout once the socket is bound. The SDKs monitor this output to capture the actual listening port before proceeding with registration.

In [`rust/src/lib.rs`](https://github.com/github/copilot-sdk/blob/main/rust/src/lib.rs), the `Client::spawn_tcp` method handles this port-wait logic, parsing the stdout stream until it detects the port announcement. Similarly, the Node.js implementation in [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts) contains the `spawnTcp` function that waits for the "TCP port wait complete" log line before returning control to your application.

### Connection Token Security

When an SDK spawns its own CLI instance, it automatically generates a **UUID token** unless you provide one explicitly. This token must be included in the JSON-RPC `connect` request, and the CLI validates it before processing subsequent commands. This mechanism prevents unauthorized processes from hijacking TCP connections in a multi-tenant environment.

As implemented in the Rust client (`Client::start` in [`rust/src/lib.rs`](https://github.com/github/copilot-sdk/blob/main/rust/src/lib.rs)) and the Node.js client (`token` field in `ClientOptions` in [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts)), token generation happens automatically when you set the connection parameter to `None` or `null`.

### Graceful Shutdown Handling

When a CLI process exits, the SDK automatically closes the TCP socket and removes the backend from the balancer's pool. This ensures clients do not receive half-open connections during scaling events, rolling deployments, or process restarts.

## Load Balancer Configuration Requirements

Proper load balancer configuration ensures optimal performance and reliability when scaling Copilot CLI.

### TCP vs HTTP Routing

Configure your load balancer to operate in **TCP mode** (Layer 4) rather than HTTP mode (Layer 7). The Copilot CLI uses raw JSON-RPC streams over TCP, not HTTP requests. HTTP-aware load balancers may buffer, modify, or terminate the stream, breaking the JSON-RPC protocol. Use raw TCP forwarding to preserve message framing.

### Health Check Strategy

Implement TCP health checks that attempt a simple JSON-RPC `ping` to verify backend health. Because the protocol is stateless, a failed health check indicates the specific instance should be removed from the pool immediately. Configure your orchestration system to spawn replacement instances dynamically using the SDK's port discovery mechanism when health checks fail.

## SDK-Specific Implementation Patterns

Each official SDK provides utilities to spawn TCP-mode CLI instances and handle port discovery.

### Rust Implementation

The Rust SDK in [`rust/src/lib.rs`](https://github.com/github/copilot-sdk/blob/main/rust/src/lib.rs) provides the `Client::spawn_tcp` method for launching CLI instances:

```rust
use copilot_sdk::Client;
use copilot_sdk::ClientOptions;

let opts = ClientOptions::default()
    .with_mode(copilot_sdk::mode::ClientMode::CopilotCli)
    .with_tcp_port(0)               // 0 → OS‑assigned port
    .with_connection_token(None);   // SDK will generate a UUID

let client = Client::new(opts).await?;
let port = client.tcp_port().expect("TCP mode must expose a port");

// Register `port` (and the autogenerated token) with your load balancer here

```

### Node.js Implementation

The Node.js client in [`nodejs/src/client.ts`](https://github.com/github/copilot-sdk/blob/main/nodejs/src/client.ts) exposes `spawnTcp` and runtime information:

```javascript
import { CopilotClient, RuntimeConnection } from '@github/copilot-sdk';

const conn = RuntimeConnection.forTcp({ port: 0 }); // OS picks a free port
const client = new CopilotClient({ runtimeConnection: conn });

await client.start();                     // waits for port announcement
const { port, token } = client.runtimeInfo; // expose to load balancer

```

### Python Implementation

Python developers use `RuntimeConnection.for_tcp` as shown in [`python/copilot/client.py`](https://github.com/github/copilot-sdk/blob/main/python/copilot/client.py) and documented in [`python/README.md`](https://github.com/github/copilot-sdk/blob/main/python/README.md):

```python
from copilot import client, runtime

# Create a TCP runtime; port=0 lets the OS pick a free port

runtime_conn = runtime.RuntimeConnection.for_tcp(port=0)

c = client.CopilotClient(runtime_connection=runtime_conn)
c.start()               # blocks until the CLI prints its port

port = c.runtime_info.port
token = c.runtime_info.connection_token

# Now add `port`/`token` to your load‑balancer pool

```

### Java Implementation

The Java SDK in [`java/src/main/java/com/github/copilot/rpc/CopilotClientOptions.java`](https://github.com/github/copilot-sdk/blob/main/java/src/main/java/com/github/copilot/rpc/CopilotClientOptions.java) provides getters and setters for TCP configuration:

```java
CopilotClientOptions opts = new CopilotClientOptions()
    .setMode(ClientMode.COPILOT_CLI)
    .setTcpPort(0)                    // let OS allocate
    .setConnectionToken(null);        // SDK will auto‑generate

CopilotClient client = new CopilotClient(opts);
client.start();                      // waits for port announcement
int port = client.getTcpPort();
String token = client.getConnectionToken();

```

## Summary

- **Stateless architecture**: Copilot CLI maintains no disk state, enabling safe horizontal scaling across multiple instances behind a load balancer.
- **TCP mode requirement**: Use TCP transport (not stdio or HTTP) when spawning CLI instances to ensure compatibility with standard load balancers.
- **Automatic port discovery**: Set port to `0` and let the SDK parse the `PORT=<num>` stdout announcement from the CLI process.
- **Security via tokens**: Auto-generated UUID tokens in the JSON-RPC `connect` request prevent unauthorized access to CLI instances.
- **Load balancer configuration**: Use Layer 4 TCP load balancing with JSON-RPC ping health checks, removing failed instances automatically when connections close.
- **Consistent configuration**: Pass `experimentAssignments`, `skillDirectories`, and `otelConfig` via `ClientOptions` to ensure uniform behavior across all instances.

## Frequently Asked Questions

### Does Copilot CLI maintain session state that affects load balancing?

No. Copilot CLI is completely stateless and does not write user data to disk. All session information resides in memory, meaning any instance can handle any request. This design eliminates the need for sticky sessions or shared session stores when configuring your load balancer.

### What load balancer algorithms work best with Copilot CLI?

Round-robin or least-connections algorithms work well because the JSON-RPC protocol is stateless and uniform. Avoid algorithms that assume HTTP semantics or require session affinity. Ensure your load balancer supports raw TCP forwarding (Layer 4) rather than HTTP proxying (Layer 7) to prevent protocol corruption.

### How do I handle CLI process failures in a load-balanced setup?

Monitor TCP connection health using JSON-RPC `ping` requests. When a CLI process exits, the SDK closes the socket automatically, causing health checks to fail. Configure your load balancer to remove failed backends immediately and spawn replacement instances dynamically using the SDK's port discovery mechanism.

### Can I use HTTP/HTTPS load balancers instead of TCP?

No. Copilot CLI uses raw JSON-RPC over TCP, not HTTP. Using an HTTP-aware load balancer will corrupt the protocol stream or cause connection failures. Always configure Layer 4 TCP forwarding to preserve the binary JSON-RPC message framing between the client and CLI instances.