How to Debug Grok Egress Connection Failures and Transport Errors
To debug Grok egress failures, inspect the structured JSON error responses from the API endpoints, classify transport errors using the neterror package helpers like IsResponseHeaderTimeout, and verify node health metrics and quality guard states via the management endpoints.
The chenyme/grok2api project implements a specialized egress layer for routing requests to Grok nodes, but transport failures can occur due to network timeouts, proxy misconfigurations, or Cloudflare clearance issues. When a node becomes unreachable, the system surfaces errors through a structured HTTP API that exposes detailed diagnostic information. Understanding how to interpret these errors requires familiarity with the internal error classification system located in backend/internal/pkg/neterror/classify.go and the specific endpoints designed for health checking.
Step-by-Step Debugging Workflow
Inspect Structured API Error Responses
All egress-related endpoints return a structured JSON payload containing an error field when operations fail. The HTTP handler in backend/internal/transport/http/egress/handler.go uses the writeError method (lines 71-78) to serialize error conditions into the response body.
When testing a specific node via POST /egress-nodes/:id/test, transport failures surface through this mechanism with specific error codes. For example, a Cloudflare clearance failure returns clearanceRefreshFailed, while network timeouts yield distinct classification strings.
To manually trigger a diagnostic check:
curl -X POST http://localhost:8080/api/egress-nodes/42/test \
-H "Content-Type: application/json" \
-d '{}' | jq .
A failed response appears as:
{
"error": "clearanceRefreshFailed",
"message": "cloudflare clearance failed",
"details": "timeout awaiting response headers"
}
Classify Errors Using the Neterror Package
The system categorizes transport errors using sentinel values defined in backend/internal/pkg/neterror/classify.go. The handler maps raw errors to these classifications via the writeError function, which checks against types like egressapp.ErrInvalidInput and egressapp.ErrProbeStale.
For programmatic detection, import the package and use the classification helpers:
import "github.com/chenyme/grok2api/backend/internal/pkg/neterror"
if neterror.IsResponseHeaderTimeout(err) {
// Handle header timeout
} else if neterror.IsUpstreamStreamIdleTimeout(err) {
// Handle stream idle
}
Distinguish Header Timeouts from Stream Idle Timeouts
Two specific timeout conditions require different remediation strategies. The neterror package provides distinct detection functions in classify.go:
IsResponseHeaderTimeout(lines 20-31): Returns true when the Go HTTP client times out waiting for the first response headers, indicating the remote server is down or extremely slow.IsUpstreamStreamIdleTimeout(lines 39-43): Returns true when a provider-side stream aborts because no data arrived within the configured idle window, typically used by the Build service.
When you encounter these errors in the service layer (backend/internal/application/egress/service.go), the methods ProbeQuality and TestNode return these specific error types before the handler wraps them for HTTP transmission.
Verify Node Health and Quality Guard Status
If the egress node is being throttled by the quality guard, the system may have disabled it automatically. Check the guard state via GET /egress-quality-guard, which reads from the JSON file specified by guardStatePath.
Look for disabledByGuard: true in the node state. Additionally, examine the Health field returned in GET /egress-nodes—a value of 0% usually indicates repeated transport errors. The newNodeResponse function in handler.go (lines 10-20) constructs these payloads, including health calculations from recent probe results.
Check Proxy and Cloudflare Configurations
Many transport failures stem from misconfigured proxies or missing Cloudflare clearance cookies. The writeError method in handler.go (lines 55-62) specifically checks for error messages containing "FlareSolverr" or "Clearance", mapping them to the clearanceRefreshFailed response with HTTP 502.
Ensure the node's proxyUrl and cloudflareCookies fields are correctly configured. If clearance refresh fails, the error details will explicitly mention timeout issues awaiting response headers from the Cloudflare challenge endpoint.
Enable Debug Logging for Transport Details
The service layer logs original errors before wrapping them. Set the environment variable LOG_LEVEL=debug when running the container to capture detailed HTTP request/response cycles, including URLs, timeout values, and intermediate redirects.
Access these logs via:
docker logs <container>
Look for messages such as "upstream stream idle timeout" or "timeout awaiting response headers", which originate from the sentinel errors in the classification layer. The logger implementation respects this level across files like backend/internal/infra/runtime/redis/store.go and other components using the standard logging interface.
Tune Transport Timeouts
If you operate slow egress endpoints, you may need to adjust the underlying http.Transport timeout settings. The HTTP client factory in backend/internal/infra/runtime/httpclient.go creates the client used by the egress service. Modify the transport configuration there to increase dial or response header timeouts for specific network conditions.
Practical Debugging Examples
The following Go example demonstrates how to detect specific timeout conditions in a custom health-check script:
package main
import (
"fmt"
"net/http"
"time"
"github.com/chenyme/grok2api/backend/internal/pkg/neterror"
)
func main() {
cli := &http.Client{Timeout: 30 * time.Second}
_, err := cli.Get("https://example.com/slow-endpoint")
if err != nil {
if neterror.IsResponseHeaderTimeout(err) {
fmt.Println("Header timeout – remote server is not responding")
} else if neterror.IsUpstreamStreamIdleTimeout(err) {
fmt.Println("Stream idle timeout – provider stopped sending data")
} else {
fmt.Printf("Other transport error: %v\n", err)
}
}
}
To read the quality guard state directly for offline debugging:
package main
import (
"encoding/json"
"fmt"
"os"
)
func main() {
// Mirrors readQualityGuardState logic from the handler
data, err := os.ReadFile("/data/guard_state.json")
if err != nil {
panic(err)
}
var state struct {
Guard struct{ Mode string }
Nodes map[string]interface{}
}
json.Unmarshal(data, &state)
fmt.Printf("Guard mode: %s, active nodes: %d\n", state.Guard.Mode, len(state.Nodes))
}
Key Source Files for Debugging
| File | Role |
|---|---|
backend/internal/transport/http/egress/handler.go |
HTTP API layer that routes requests, parses inputs, and maps errors to JSON responses via writeError. |
backend/internal/pkg/neterror/classify.go |
Centralized error classification containing IsResponseHeaderTimeout and IsUpstreamStreamIdleTimeout helpers. |
backend/internal/application/egress/service.go |
Business logic that communicates with remote egress providers, performs probes, and returns raw transport errors. |
backend/internal/domain/egress/node.go |
Domain model defining the egress node structure, including health metrics and guard-related fields. |
backend/internal/infra/runtime/httpclient.go |
Factory for http.Client instances used by the egress service, configuring timeouts and proxy settings. |
Summary
- Inspect API responses from endpoints like
POST /egress-nodes/:id/testto obtain structured error codes such asclearanceRefreshFailed. - Use
neterrorhelpers (IsResponseHeaderTimeout,IsUpstreamStreamIdleTimeout) to classify whether failures stem from initial connection issues or mid-stream data interruptions. - Check quality guard states via
GET /egress-quality-guardto determine if nodes are disabled due to repeated failures. - Verify proxy and Cloudflare configurations when encountering 502 errors related to clearance refresh failures.
- Enable debug logging with
LOG_LEVEL=debugto view raw transport cycles and correlate errors with specific HTTP phases.
Frequently Asked Questions
What does the "clearanceRefreshFailed" error indicate?
The clearanceRefreshFailed error indicates that the system cannot obtain valid Cloudflare clearance cookies for the egress node. According to the handler implementation in backend/internal/transport/http/egress/handler.go, this occurs when the writeError function detects "FlareSolverr" or "Clearance" strings in the error message, typically caused by proxy misconfigurations or the remote server blocking automated challenges. Verify that the node's proxyUrl and cloudflareCookies fields are correctly populated.
How can I distinguish between a network timeout and a stream idle timeout?
Use the classification functions in backend/internal/pkg/neterror/classify.go. Call neterror.IsResponseHeaderTimeout(err) to detect when the Go client times out waiting for initial headers (indicating server unavailability), or neterror.IsUpstreamStreamIdleTimeout(err) to identify when an established stream stops sending data within the configured idle window. These distinctions help determine whether to investigate network connectivity versus provider-side streaming issues.
Where can I find the quality guard state for a disabled node?
Query the GET /egress-quality-guard endpoint to retrieve the current guard state, or inspect the JSON file specified by guardStatePath directly. The response contains node entries with a disabledByGuard boolean field. When true, the quality guard has automatically disabled the node due to failing health probes, which you can correlate with the Health percentage shown in GET /egress-nodes listings.
How do I increase timeout values for slow egress endpoints?
Modify the http.Transport configuration in backend/internal/infra/runtime/httpclient.go, which serves as the factory for HTTP clients used by the egress service. Adjust the dialer timeout, TLS handshake timeout, or response header timeout values to accommodate slow networks. After modifying these values, redeploy the container and verify the changes by running manual probe tests against the affected nodes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →