# How to Debug Grok Egress Connection Failures and Transport Errors

> Debug Grok egress connection failures and transport errors by inspecting API JSON, using neterror helpers, and verifying node health metrics. Resolve your Grok issues efficiently.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: how-to-guide
- Published: 2026-08-09

---

**To debug Grok egress failures, inspect the structured JSON error responses from the API endpoints, classify transport errors using the `neterror` package helpers like `IsResponseHeaderTimeout`, and verify node health metrics and quality guard states via the management endpoints.**

The `chenyme/grok2api` project implements a specialized egress layer for routing requests to Grok nodes, but transport failures can occur due to network timeouts, proxy misconfigurations, or Cloudflare clearance issues. When a node becomes unreachable, the system surfaces errors through a structured HTTP API that exposes detailed diagnostic information. Understanding how to interpret these errors requires familiarity with the internal error classification system located in [`backend/internal/pkg/neterror/classify.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/pkg/neterror/classify.go) and the specific endpoints designed for health checking.

## Step-by-Step Debugging Workflow

### Inspect Structured API Error Responses

All egress-related endpoints return a structured JSON payload containing an `error` field when operations fail. The HTTP handler in [`backend/internal/transport/http/egress/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/egress/handler.go) uses the `writeError` method (lines 71-78) to serialize error conditions into the response body.

When testing a specific node via `POST /egress-nodes/:id/test`, transport failures surface through this mechanism with specific error codes. For example, a Cloudflare clearance failure returns `clearanceRefreshFailed`, while network timeouts yield distinct classification strings.

To manually trigger a diagnostic check:

```bash
curl -X POST http://localhost:8080/api/egress-nodes/42/test \
     -H "Content-Type: application/json" \
     -d '{}' | jq .

```

A failed response appears as:

```json
{
  "error": "clearanceRefreshFailed",
  "message": "cloudflare clearance failed",
  "details": "timeout awaiting response headers"
}

```

### Classify Errors Using the Neterror Package

The system categorizes transport errors using sentinel values defined in [`backend/internal/pkg/neterror/classify.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/pkg/neterror/classify.go). The handler maps raw errors to these classifications via the `writeError` function, which checks against types like `egressapp.ErrInvalidInput` and `egressapp.ErrProbeStale`.

For programmatic detection, import the package and use the classification helpers:

```go
import "github.com/chenyme/grok2api/backend/internal/pkg/neterror"

if neterror.IsResponseHeaderTimeout(err) {
    // Handle header timeout
} else if neterror.IsUpstreamStreamIdleTimeout(err) {
    // Handle stream idle
}

```

### Distinguish Header Timeouts from Stream Idle Timeouts

Two specific timeout conditions require different remediation strategies. The `neterror` package provides distinct detection functions in [`classify.go`](https://github.com/chenyme/grok2api/blob/main/classify.go):

- **`IsResponseHeaderTimeout`** (lines 20-31): Returns true when the Go HTTP client times out waiting for the first response headers, indicating the remote server is down or extremely slow.
- **`IsUpstreamStreamIdleTimeout`** (lines 39-43): Returns true when a provider-side stream aborts because no data arrived within the configured idle window, typically used by the Build service.

When you encounter these errors in the service layer ([`backend/internal/application/egress/service.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/egress/service.go)), the methods `ProbeQuality` and `TestNode` return these specific error types before the handler wraps them for HTTP transmission.

### Verify Node Health and Quality Guard Status

If the egress node is being throttled by the quality guard, the system may have disabled it automatically. Check the guard state via `GET /egress-quality-guard`, which reads from the JSON file specified by `guardStatePath`.

Look for `disabledByGuard: true` in the node state. Additionally, examine the `Health` field returned in `GET /egress-nodes`—a value of 0% usually indicates repeated transport errors. The `newNodeResponse` function in [`handler.go`](https://github.com/chenyme/grok2api/blob/main/handler.go) (lines 10-20) constructs these payloads, including health calculations from recent probe results.

### Check Proxy and Cloudflare Configurations

Many transport failures stem from misconfigured proxies or missing Cloudflare clearance cookies. The `writeError` method in [`handler.go`](https://github.com/chenyme/grok2api/blob/main/handler.go) (lines 55-62) specifically checks for error messages containing "FlareSolverr" or "Clearance", mapping them to the `clearanceRefreshFailed` response with HTTP 502.

Ensure the node's `proxyUrl` and `cloudflareCookies` fields are correctly configured. If clearance refresh fails, the error details will explicitly mention timeout issues awaiting response headers from the Cloudflare challenge endpoint.

### Enable Debug Logging for Transport Details

The service layer logs original errors before wrapping them. Set the environment variable `LOG_LEVEL=debug` when running the container to capture detailed HTTP request/response cycles, including URLs, timeout values, and intermediate redirects.

Access these logs via:

```bash
docker logs <container>

```

Look for messages such as "upstream stream idle timeout" or "timeout awaiting response headers", which originate from the sentinel errors in the classification layer. The logger implementation respects this level across files like [`backend/internal/infra/runtime/redis/store.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/redis/store.go) and other components using the standard logging interface.

### Tune Transport Timeouts

If you operate slow egress endpoints, you may need to adjust the underlying `http.Transport` timeout settings. The HTTP client factory in [`backend/internal/infra/runtime/httpclient.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/httpclient.go) creates the client used by the egress service. Modify the transport configuration there to increase dial or response header timeouts for specific network conditions.

## Practical Debugging Examples

The following Go example demonstrates how to detect specific timeout conditions in a custom health-check script:

```go
package main

import (
	"fmt"
	"net/http"
	"time"

	"github.com/chenyme/grok2api/backend/internal/pkg/neterror"
)

func main() {
	cli := &http.Client{Timeout: 30 * time.Second}
	_, err := cli.Get("https://example.com/slow-endpoint")
	if err != nil {
		if neterror.IsResponseHeaderTimeout(err) {
			fmt.Println("Header timeout – remote server is not responding")
		} else if neterror.IsUpstreamStreamIdleTimeout(err) {
			fmt.Println("Stream idle timeout – provider stopped sending data")
		} else {
			fmt.Printf("Other transport error: %v\n", err)
		}
	}
}

```

To read the quality guard state directly for offline debugging:

```go
package main

import (
	"encoding/json"
	"fmt"
	"os"
)

func main() {
	// Mirrors readQualityGuardState logic from the handler
	data, err := os.ReadFile("/data/guard_state.json")
	if err != nil {
		panic(err)
	}
	var state struct {
		Guard struct{ Mode string }
		Nodes map[string]interface{}
	}
	json.Unmarshal(data, &state)
	fmt.Printf("Guard mode: %s, active nodes: %d\n", state.Guard.Mode, len(state.Nodes))
}

```

## Key Source Files for Debugging

| File | Role |
|------|------|
| [`backend/internal/transport/http/egress/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/egress/handler.go) | HTTP API layer that routes requests, parses inputs, and maps errors to JSON responses via `writeError`. |
| [`backend/internal/pkg/neterror/classify.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/pkg/neterror/classify.go) | Centralized error classification containing `IsResponseHeaderTimeout` and `IsUpstreamStreamIdleTimeout` helpers. |
| [`backend/internal/application/egress/service.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/application/egress/service.go) | Business logic that communicates with remote egress providers, performs probes, and returns raw transport errors. |
| [`backend/internal/domain/egress/node.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/domain/egress/node.go) | Domain model defining the egress node structure, including health metrics and guard-related fields. |
| [`backend/internal/infra/runtime/httpclient.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/httpclient.go) | Factory for `http.Client` instances used by the egress service, configuring timeouts and proxy settings. |

## Summary

- **Inspect API responses** from endpoints like `POST /egress-nodes/:id/test` to obtain structured error codes such as `clearanceRefreshFailed`.
- **Use `neterror` helpers** (`IsResponseHeaderTimeout`, `IsUpstreamStreamIdleTimeout`) to classify whether failures stem from initial connection issues or mid-stream data interruptions.
- **Check quality guard states** via `GET /egress-quality-guard` to determine if nodes are disabled due to repeated failures.
- **Verify proxy and Cloudflare configurations** when encountering 502 errors related to clearance refresh failures.
- **Enable debug logging** with `LOG_LEVEL=debug` to view raw transport cycles and correlate errors with specific HTTP phases.

## Frequently Asked Questions

### What does the "clearanceRefreshFailed" error indicate?

The `clearanceRefreshFailed` error indicates that the system cannot obtain valid Cloudflare clearance cookies for the egress node. According to the handler implementation in [`backend/internal/transport/http/egress/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/egress/handler.go), this occurs when the `writeError` function detects "FlareSolverr" or "Clearance" strings in the error message, typically caused by proxy misconfigurations or the remote server blocking automated challenges. Verify that the node's `proxyUrl` and `cloudflareCookies` fields are correctly populated.

### How can I distinguish between a network timeout and a stream idle timeout?

Use the classification functions in [`backend/internal/pkg/neterror/classify.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/pkg/neterror/classify.go). Call `neterror.IsResponseHeaderTimeout(err)` to detect when the Go client times out waiting for initial headers (indicating server unavailability), or `neterror.IsUpstreamStreamIdleTimeout(err)` to identify when an established stream stops sending data within the configured idle window. These distinctions help determine whether to investigate network connectivity versus provider-side streaming issues.

### Where can I find the quality guard state for a disabled node?

Query the `GET /egress-quality-guard` endpoint to retrieve the current guard state, or inspect the JSON file specified by `guardStatePath` directly. The response contains node entries with a `disabledByGuard` boolean field. When `true`, the quality guard has automatically disabled the node due to failing health probes, which you can correlate with the `Health` percentage shown in `GET /egress-nodes` listings.

### How do I increase timeout values for slow egress endpoints?

Modify the `http.Transport` configuration in [`backend/internal/infra/runtime/httpclient.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/httpclient.go), which serves as the factory for HTTP clients used by the egress service. Adjust the dialer timeout, TLS handshake timeout, or response header timeout values to accommodate slow networks. After modifying these values, redeploy the container and verify the changes by running manual probe tests against the affected nodes.