How to Configure Egress Proxy Pools with Fallback and Balancing in Grok2API
Grok2API routes outbound requests through configurable egress nodes that support proxy pools, fixed fallback policies, and automatic load balancing to prevent cascade failures and optimize traffic distribution.
The chenyme/grok2api repository implements a sophisticated egress management system that treats outbound traffic through three distinct resilience mechanisms. Understanding how to configure egress proxy pools with fallback and balancing ensures high availability for Build, Web, and Console scoped requests while preventing single points of failure.
Core Mechanisms for Resilient Egress Routing
Proxy-Pool Mode
Proxy-pool mode marks an egress node as part of a shared pool of proxies rather than a single endpoint. When you enable proxyPool: true on a node, the system treats connection failures differently: only the specific failed connection is marked unusable, while the node remains available for concurrent requests. This prevents a "single-failure cascade" that would otherwise trigger cooldown for the entire node.
In backend/internal/infra/egress/manager.go, the isProxyPoolNode function identifies these nodes during request routing【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L83-L85】. The implementation ensures that requests to pool members open fresh CONNECT tunnels (freshTunnel) to avoid pinning subsequent traffic to a single unhealthy pool member【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L12-L15】.
Fixed-Fallback Policy
The fixed-fallback policy defines a stable "last-chance" node for each egress scope (Build, Web, Console, etc.). When primary node selection fails—either through health checks or exhaustion of retry attempts—the manager automatically routes requests to the configured fallback. Critically, fallback nodes must be enabled, non-pool nodes to guarantee stability.
Configuration parsing occurs in backend/internal/infra/egress/manager.go within the fixed_fallback_node_ids logic【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L294-L300】. The quality guard in tools/egress-quality-guard/quality_guard.py protects these fixed-fallback nodes from automatic quarantine, ensuring they remain available during degraded conditions【/cache/repos/github.com/chenyme/grok2api/main/tools/egress-quality-guard/quality_guard.py#L294-L300】.
Auto-Assignment and Load Balancing
Auto-assignment (autoAssignEnabled) and auto-balancing (autoBalanceEnabled) work together to distribute account capacity across nodes. Auto-assignment automatically adds newly discovered accounts to the best-available egress node for a given scope, while auto-balancing periodically redistributes accounts from overloaded nodes to under-utilized ones based on capacity metrics.
These features rely on the Egress Operations API and execute through background jobs running every assignmentIntervalSeconds. The frontend state is defined in frontend/src/features/settings/egress-operations.tsx【/cache/repos/github.com/chenyme/grok2api/main/frontend/src/features/settings/egress-operations.tsx#L58-L61】, while the backend applies these configurations through backend/internal/application/egress/service.go.
Request Routing and Failure Handling
Node Selection and Tunnel Creation
When a client request arrives, the egress manager selects an enabled node matching the request's scope. For proxy-pool nodes, the manager explicitly opens a fresh tunnel for each request to prevent session pinning to potentially degraded pool members. This ensures that each connection attempt has an independent chance of success.
Intelligent Retry Logic
If a request fails with a connection error and the target node is a proxy-pool member, the manager implements retry logic with a proxyPoolRetryLimit of 2 attempts. The system routes the retried request through a different pool member, effectively bypassing the failed proxy without marking the entire node as unhealthy【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/infra/egress/manager.go#L36-L38】.
Fallback Escalation
When all primary candidates are exhausted, the manager escalates to the fixed-fallback node configured for the request's specific scope. This escalation happens automatically if the fallbacks configuration maps the current scope to a valid node ID. The fallback mechanism ensures that even during widespread node degradation, traffic can still egress through the protected fallback path.
Configuration Implementation
Enabling Proxy-Pool Mode
To configure a node as part of a proxy pool, enable the proxy-pool switch in the node editor. The UI component in frontend/src/features/settings/egress-nodes.tsx controls this setting:
// frontend/src/features/settings/egress-nodes.tsx (excerpt)
<Switch
id="egress-proxy-pool"
checked={form.proxyPool}
disabled={!editing?.proxyConfigured && !form.proxyURL?.trim()}
onCheckedChange={(proxyPool) => setForm({ ...form, proxyPool })}
/>
When saved, the backend handler in backend/internal/transport/http/egress/handler.go processes the proxyPool JSON field to persist the configuration【/cache/repos/github.com/chenyme/grok2api/main/backend/internal/transport/http/egress/handler.go#L418-L433】.
Configuring Fixed Fallbacks
Define fallback policies through the Egress Operations API. The following JSON payload configures a fixed fallback for the Build scope while disabling fallbacks for other scopes:
{
"probeProvider": "cloudflare",
"probeIntervalSeconds": 900,
"autoAssignEnabled": true,
"autoBalanceEnabled": true,
"assignmentIntervalSeconds": 300,
"fallbacks": {
"grok_build": { "mode": "fixed", "nodeId": "12" },
"grok_web": { "mode": "none" },
"grok_console": { "mode": "none" },
"grok_web_asset": { "mode": "none" },
"grok_console_asset": { "mode": "none" }
},
"subscriptionProxyURL": "",
"subscriptionProxyConfigured": false
}
Send this payload to POST /api/admin/v1/egress-operations, handled by updateEgressOperationsConfig in frontend/src/features/settings/settings-api.ts.
Activating Auto-Balance and Auto-Assign
Enable automatic load distribution through the operations panel. The React component manages these boolean flags:
// frontend/src/features/settings/egress-operations.tsx (excerpt)
<Switch
id="egress-auto-assign"
checked={operationsForm.autoAssignEnabled}
onCheckedChange={(v) =>
setOperationsDraft({ ...operationsForm, autoAssignEnabled: v })
}
/>
<Switch
id="egress-auto-balance"
checked={operationsForm.autoBalanceEnabled}
onCheckedChange={(v) =>
setOperationsDraft({ ...operationsForm, autoBalanceEnabled: v })
}
/>
When enabled, the system periodically invokes rebalanceEgressAccounts to redistribute accounts based on node health and capacity metrics.
Summary
- Proxy-pool mode isolates connection failures to individual pool members rather than marking entire nodes unhealthy, implemented through
isProxyPoolNodeinmanager.go. - Fixed-fallback policies provide guaranteed egress paths when primary nodes fail, with protection against quarantine in
quality_guard.py. - Auto-assign and auto-balance continuously optimize account distribution across nodes using background jobs configured via
assignmentIntervalSeconds. - Retry logic attempts up to 2 alternate pool members before escalating to fallback nodes.
- Configuration occurs through the Egress Operations API and Egress Nodes UI, with schemas defined in
handler.goandsettings-api.ts.
Frequently Asked Questions
What is the difference between a proxy-pool node and a standard egress node?
A standard egress node represents a single proxy endpoint or direct connection, where any connection failure typically triggers cooldown for the entire node. A proxy-pool node (proxyPool: true) treats each connection attempt independently—if one pool member fails, only that specific connection is marked failed, and the node remains available for other requests. This prevents cascading failures while allowing retry attempts on alternate pool members.
How does the fixed-fallback mechanism protect against total egress failure?
The fixed-fallback mechanism ensures that when all primary nodes for a scope (Build, Web, Console) are unhealthy or exhausted, traffic automatically routes to a pre-configured stable node. These fallback nodes are explicitly protected from automatic quarantine by the quality guard (fixed_fallback_node_ids in quality_guard.py), guaranteeing a last-resort egress path even during widespread degradation.
What triggers the auto-balance functionality, and how often does it run?
Auto-balance triggers when autoBalanceEnabled is set to true in the Egress Operations configuration. The background job runs every assignmentIntervalSeconds (default typically 300 seconds), redistributing accounts from overloaded nodes to under-utilized ones based on current capacity metrics. This runs independently from auto-assign, which only handles newly discovered accounts.
Why does the system use fresh tunnels for proxy-pool requests?
The system opens fresh tunnels (freshTunnel) for proxy-pool requests to avoid HTTP connection reuse (keep-alive) that would pin subsequent requests to the same potentially degraded pool member. This ensures each request has an independent connection attempt, maximizing the probability of success when routing through a pool of proxies with varying health states.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →