# How to Configure Egress Proxy Pools with Fallback and Balancing in Grok2API

> Configure egress proxy pools with fallback and balancing in Grok2API. Optimize traffic distribution and prevent cascade failures with smart routing policies.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Grok2API routes outbound requests through configurable egress nodes that support proxy pools, fixed fallback policies, and automatic load balancing to prevent cascade failures and optimize traffic distribution.**

The `chenyme/grok2api` repository implements a sophisticated egress management system that treats outbound traffic through three distinct resilience mechanisms. Understanding how to configure egress proxy pools with fallback and balancing ensures high availability for Build, Web, and Console scoped requests while preventing single points of failure.

## Core Mechanisms for Resilient Egress Routing

### Proxy-Pool Mode

Proxy-pool mode marks an egress node as part of a shared pool of proxies rather than a single endpoint. When you enable `proxyPool: true` on a node, the system treats connection failures differently: only the specific failed connection is marked unusable, while the node remains available for concurrent requests. This prevents a "single-failure cascade" that would otherwise trigger cooldown for the entire node.

In [`backend/internal/infra/egress/manager.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/manager.go), the `isProxyPoolNode` function identifies these nodes during request routing【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L83-L85】. The implementation ensures that requests to pool members open fresh CONNECT tunnels (`freshTunnel`) to avoid pinning subsequent traffic to a single unhealthy pool member【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L12-L15】.

### Fixed-Fallback Policy

The fixed-fallback policy defines a stable "last-chance" node for each egress scope (Build, Web, Console, etc.). When primary node selection fails—either through health checks or exhaustion of retry attempts—the manager automatically routes requests to the configured fallback. Critically, fallback nodes must be **enabled, non-pool** nodes to guarantee stability.

Configuration parsing occurs in [`backend/internal/infra/egress/manager.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/egress/manager.go) within the `fixed_fallback_node_ids` logic【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L294-L300】. The quality guard in [`tools/egress-quality-guard/quality_guard.py`](https://github.com/chenyme/grok2api/blob/main/tools/egress-quality-guard/quality_guard.py) protects these fixed-fallback nodes from automatic quarantine, ensuring they remain available during degraded conditions【/cache/repos/github.com/chenyme/grok2api/main/tools/egress-quality-guard/quality_guard.py#L294-L300】.

### Auto-Assignment and Load Balancing

Auto-assignment (`autoAssignEnabled`) and auto-balancing (`autoBalanceEnabled`) work together to distribute account capacity across nodes. Auto-assignment automatically adds newly discovered accounts to the best-available egress node for a given scope, while auto-balancing periodically redistributes accounts from overloaded nodes to under-utilized ones based on capacity metrics.

These features rely on the **Egress Operations** API and execute through background jobs running every `assignmentIntervalSeconds`. The frontend state is defined in [`frontend/src/features/settings/egress-operations.tsx`](https://github.com/chenyme/grok2api/blob/main/frontend/src/features/settings/egress-operations.tsx)【/cache/repos/github.com/chenyme/grok2api/main/frontend/src/features/settings/egress-operations.tsx#L58-L61】, while the backend applies these configurations through [`backend/internal/application/egress/service.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/egress/service.go).

## Request Routing and Failure Handling

### Node Selection and Tunnel Creation

When a client request arrives, the egress manager selects an enabled node matching the request's scope. For proxy-pool nodes, the manager explicitly opens a fresh tunnel for each request to prevent session pinning to potentially degraded pool members. This ensures that each connection attempt has an independent chance of success.

### Intelligent Retry Logic

If a request fails with a connection error and the target node is a proxy-pool member, the manager implements retry logic with a `proxyPoolRetryLimit` of 2 attempts. The system routes the retried request through a different pool member, effectively bypassing the failed proxy without marking the entire node as unhealthy【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L36-L38】.

### Fallback Escalation

When all primary candidates are exhausted, the manager escalates to the fixed-fallback node configured for the request's specific scope. This escalation happens automatically if the `fallbacks` configuration maps the current scope to a valid node ID. The fallback mechanism ensures that even during widespread node degradation, traffic can still egress through the protected fallback path.

## Configuration Implementation

### Enabling Proxy-Pool Mode

To configure a node as part of a proxy pool, enable the proxy-pool switch in the node editor. The UI component in [`frontend/src/features/settings/egress-nodes.tsx`](https://github.com/chenyme/grok2api/blob/main/frontend/src/features/settings/egress-nodes.tsx) controls this setting:

```tsx
// frontend/src/features/settings/egress-nodes.tsx (excerpt)
<Switch
  id="egress-proxy-pool"
  checked={form.proxyPool}
  disabled={!editing?.proxyConfigured && !form.proxyURL?.trim()}
  onCheckedChange={(proxyPool) => setForm({ ...form, proxyPool })}
/>

```

When saved, the backend handler in [`backend/internal/transport/http/egress/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/egress/handler.go) processes the `proxyPool` JSON field to persist the configuration【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/transport/http/egress/handler.go#L418-L433】.

### Configuring Fixed Fallbacks

Define fallback policies through the Egress Operations API. The following JSON payload configures a fixed fallback for the Build scope while disabling fallbacks for other scopes:

```json
{
  "probeProvider": "cloudflare",
  "probeIntervalSeconds": 900,
  "autoAssignEnabled": true,
  "autoBalanceEnabled": true,
  "assignmentIntervalSeconds": 300,
  "fallbacks": {
    "grok_build":   { "mode": "fixed", "nodeId": "12" },
    "grok_web":     { "mode": "none" },
    "grok_console": { "mode": "none" },
    "grok_web_asset": { "mode": "none" },
    "grok_console_asset": { "mode": "none" }
  },
  "subscriptionProxyURL": "",
  "subscriptionProxyConfigured": false
}

```

Send this payload to `POST /api/admin/v1/egress-operations`, handled by `updateEgressOperationsConfig` in [`frontend/src/features/settings/settings-api.ts`](https://github.com/chenyme/grok2api/blob/main/frontend/src/features/settings/settings-api.ts).

### Activating Auto-Balance and Auto-Assign

Enable automatic load distribution through the operations panel. The React component manages these boolean flags:

```tsx
// frontend/src/features/settings/egress-operations.tsx (excerpt)
<Switch
  id="egress-auto-assign"
  checked={operationsForm.autoAssignEnabled}
  onCheckedChange={(v) =>
    setOperationsDraft({ ...operationsForm, autoAssignEnabled: v })
  }
/>
<Switch
  id="egress-auto-balance"
  checked={operationsForm.autoBalanceEnabled}
  onCheckedChange={(v) =>
    setOperationsDraft({ ...operationsForm, autoBalanceEnabled: v })
  }
/>

```

When enabled, the system periodically invokes `rebalanceEgressAccounts` to redistribute accounts based on node health and capacity metrics.

## Summary

- **Proxy-pool mode** isolates connection failures to individual pool members rather than marking entire nodes unhealthy, implemented through `isProxyPoolNode` in [`manager.go`](https://github.com/chenyme/grok2api/blob/main/manager.go).
- **Fixed-fallback policies** provide guaranteed egress paths when primary nodes fail, with protection against quarantine in [`quality_guard.py`](https://github.com/chenyme/grok2api/blob/main/quality_guard.py).
- **Auto-assign and auto-balance** continuously optimize account distribution across nodes using background jobs configured via `assignmentIntervalSeconds`.
- **Retry logic** attempts up to 2 alternate pool members before escalating to fallback nodes.
- Configuration occurs through the **Egress Operations** API and **Egress Nodes** UI, with schemas defined in [`handler.go`](https://github.com/chenyme/grok2api/blob/main/handler.go) and [`settings-api.ts`](https://github.com/chenyme/grok2api/blob/main/settings-api.ts).

## Frequently Asked Questions

### What is the difference between a proxy-pool node and a standard egress node?

A standard egress node represents a single proxy endpoint or direct connection, where any connection failure typically triggers cooldown for the entire node. A **proxy-pool node** (`proxyPool: true`) treats each connection attempt independently—if one pool member fails, only that specific connection is marked failed, and the node remains available for other requests. This prevents cascading failures while allowing retry attempts on alternate pool members.

### How does the fixed-fallback mechanism protect against total egress failure?

The fixed-fallback mechanism ensures that when all primary nodes for a scope (Build, Web, Console) are unhealthy or exhausted, traffic automatically routes to a pre-configured stable node. These fallback nodes are explicitly protected from automatic quarantine by the quality guard (`fixed_fallback_node_ids` in [`quality_guard.py`](https://github.com/chenyme/grok2api/blob/main/quality_guard.py)), guaranteeing a last-resort egress path even during widespread degradation.

### What triggers the auto-balance functionality, and how often does it run?

Auto-balance triggers when `autoBalanceEnabled` is set to true in the Egress Operations configuration. The background job runs every `assignmentIntervalSeconds` (default typically 300 seconds), redistributing accounts from overloaded nodes to under-utilized ones based on current capacity metrics. This runs independently from auto-assign, which only handles newly discovered accounts.

### Why does the system use fresh tunnels for proxy-pool requests?

The system opens **fresh tunnels** (`freshTunnel`) for proxy-pool requests to avoid HTTP connection reuse (keep-alive) that would pin subsequent requests to the same potentially degraded pool member. This ensures each request has an independent connection attempt, maximizing the probability of success when routing through a pool of proxies with varying health states.