# How to Debug MasterDnsVPN Resolver Health Check Failures and Auto-Disable Behavior

> Fix MasterDnsVPN resolver health check failures. Learn to debug unexpected dropouts by inspecting connectionStats and verifying the autoDisableMinObservationsForActiveCount threshold in balancer.go.

- Repository: [Amin Mahmoudi/MasterDnsVPN](https://github.com/masterking32/MasterDnsVPN)
- Tags: how-to-guide
- Published: 2026-05-10

---

**MasterDnsVPN disables DNS resolvers when they reach a 100% loss rate inside a configurable sliding window, but you can debug unexpected dropouts by inspecting `connectionStats` in [`internal/client/balancer.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/balancer.go) and verifying the `autoDisableMinObservationsForActiveCount` threshold.**

MasterDnsVPN continuously probes its upstream DNS resolvers to ensure they remain reachable and performant. When a resolver repeatedly fails to answer a probe, the balancer marks it as **invalid** and optionally **auto-disables** it to prevent traffic from routing to a dead endpoint. Understanding the three-part health-check architecture—probe scheduling, timeout aggregation, and background re-checking—is essential for diagnosing why resolvers disappear from the active set.

## How the Health-Check System Works

The health-check system is split across three distinct components that track probe lifecycle from transmission to auto-disable decision.

### Probe Scheduling and Tracking

Every outbound DNS query triggers `TrackResolverSend` in [`internal/client/balancer.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/balancer.go). This method stores a `balancerResolverSample` keyed by resolver address, local address, and DNS-ID, recording the exact send time. If a response arrives, `TrackResolverSuccess` removes the sample and retracts any previously recorded timeout via `RetractTimeout`. If no response arrives before the deadline, `TrackResolverFailure` invokes `ReportTimeout`, which increments the **lost** counter in the resolver’s `connectionStats` structure.

### Timeout Aggregation and Auto-Disable Logic

`ReportTimeout` aggregates failures into a sliding-window statistic. It checks whether the resolver has sent enough probes to meet the minimum observation count (`autoDisableMinObservationsForActiveCount`) and whether **all** probes within the current window have timed out. When both conditions are met, the code may invoke an optional `confirmResolverDown` handler. If that handler returns true (or is nil), the balancer marks the resolver invalid via `SetConnectionValidityWithLog` and triggers `onResolverDisabled`.

### Background Health Checks for Inactive Resolvers

When `cfg.RecheckInactiveServersEnabled` is true, `runResolverHealthLoop` in [`internal/client/mtu.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/mtu.go) periodically selects inactive resolvers using `NextInactiveConnectionForHealthCheck`. It runs a full MTU probe (`recheckInactiveResolver`) against each candidate. A successful MTU test re-enables the resolver by calling `SetConnectionValidityWithLog(true, true)`, while failures keep it disabled until the next iteration.

## Common Reasons for Unexpected Auto-Disable

Several configuration and environmental factors can cause premature or confusing resolver disablement:

- **`autoDisableEnabled` is false** – When disabled, the balancer records loss statistics but never calls `SetConnectionValidityWithLog` to disable the resolver. Verify the flag via `balancer.autoDisableEnabled` or by checking the startup logs for the `SetAutoDisableConfig` initialization.

- **Low active resolver count** – The function `autoDisableMinObservationsForActiveCount(active, window)` returns a very low threshold (often 1) when fewer than 3 resolvers are active. This causes immediate disablement after a single timeout rather than tolerating transient packet loss.

- **Missing confirmation handler** – If `confirmResolverDown` is nil, the balancer disables the resolver immediately upon hitting the 100% loss threshold without secondary verification. Set a custom handler with `SetResolverDownConfirmHandler` to inject an extra probe or logging step.

- **Disabled re-check loop** – Inactive resolvers never re-enable if `RecheckInactiveServersEnabled` is false or if `resolverHealthRecheckInterval` is set too conservatively. Ensure the background health loop is running and the poll interval (`resolverHealthPollInterval`) is sufficiently aggressive.

## Step-by-Step Debugging Procedures

Follow these steps to isolate the exact failure point in the health-check pipeline:

1. **Enable debug-level logging**  
   Set the balancer logger to `DEBUG` to capture per-probe outcomes:
   ```go
   balancer.log.SetLevel(logger.LevelDebug)
   // Or via CLI: --log-level=debug
   ```

2. **Inspect sliding-window statistics**  
   Dump the raw counters for a specific resolver key to see sent vs. lost ratios:
   ```go
   stats := balancer.statsForKey("resolver-key-here")
   fmt.Printf("sent=%d lost=%d windowSent=%d windowLost=%d\n",
       stats.sent.Load(), stats.lost.Load(),
       stats.windowSent.Load(), stats.windowLost.Load())
   ```

3. **Force a rapid timeout**  
   Temporarily reduce `resolverHealthProbeTimeout` to 200 ms and enable auto-disable with a short window to reproduce the failure quickly:
   ```go
   balancer.SetAutoDisableConfig(true, 2*time.Second)
   ```

4. **Monitor pending-sample sweeps**  
   Add instrumentation at the start of `CollectExpiredResolverTimeouts` to see how many stale probes accumulate. High pending counts indicate the success path is failing to clean up entries.

5. **Implement a confirmation logger**  
   Inject a confirmation handler to audit the disable decision:
   ```go
   balancer.SetResolverDownConfirmHandler(func(c *client.Connection, w time.Duration) bool {
       fmt.Printf("confirm down for %s (window %v)\n", c.Key, w)
       return true // return false to block disablement
   })
   ```

6. **Trigger a manual re-check**  
   Test an inactive resolver directly to verify the re-enable path:
   ```go
   conn, _ := balancer.GetConnectionByKey("resolver-key")
   client.recheckInactiveResolver(context.Background(), conn)
   ```

   Look for log lines containing `✅ Accepted` (MTU success) followed by `🟢` balancer messages indicating the resolver returned to the active pool.

## Practical Example: Diagnosing a Disabled Resolver

Use this debug function immediately after observing a "DNS Resolver disabled" warning to determine whether the resolver met the all-timeouts rule or was forced by a custom handler:

```go
func debugResolver(key string, b *client.Balancer) {
    // 1. Verify auto-disable is active
    fmt.Printf("Auto-disable enabled: %v\n", b.autoDisableEnabled)

    // 2. Calculate the observation window for current active count
    active := b.ActiveCount()
    window := time.Duration(b.autoDisableTimeoutWindow) * time.Second
    fmt.Printf("Active resolvers: %d, window: %v\n", active, window)

    // 3. Dump loss statistics
    if stats := b.statsForKey(key); stats != nil {
        fmt.Printf("sent=%d lost=%d\n", stats.sent.Load(), stats.lost.Load())
    }

    // 4. Check pending sample queue depth
    fmt.Printf("Pending samples: %d\n", b.pendingCount())
}

```

Call `debugResolver("<resolver-key>", balancer)` to see if the `lost` count equals the `sent` count within the calculated window, confirming the 100% loss trigger fired.

## Critical Configuration Parameters

These methods and flags control sensitivity and recovery behavior:

- **`SetAutoDisableConfig(enabled bool, window time.Duration)`** – Enables auto-disable and sets the sliding-window length (typically `5s` to `30s`) over which loss is measured.

- **`SetResolverDisabledHandler(fn func(*Connection, string))`** – Invoked when a resolver is marked invalid. Use this to emit metrics or pager alerts.

- **`SetResolverDownConfirmHandler(fn func(*Connection, time.Duration) bool)`** – Optional gatekeeper that executes before disablement. Return `true` to allow the disable, `false` to keep the resolver active.

- **`cfg.RecheckInactiveServersEnabled`** – Boolean toggle for the background health loop in [`internal/client/mtu.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/mtu.go) that attempts to revive failed resolvers.

- **`resolverHealthRecheckInterval` / `resolverHealthPollInterval`** – Control how frequently inactive resolvers are re-probed and how long the health loop sleeps between sweeps.

## Key Source Files

| File | Purpose |
|------|---------|
| [`internal/client/balancer.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/balancer.go) | Core balancing, probe tracking via `TrackResolverSend`, timeout aggregation in `ReportTimeout`, and auto-disable logic. |
| [`internal/client/mtu.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/mtu.go) | Background health loop (`runResolverHealthLoop`), inactive resolver selection (`NextInactiveConnectionForHealthCheck`), and re-enablement via MTU probes. |
| [`internal/client/ping_manager.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/ping_manager.go) | Controls ping intervals; high-frequency pings can mask resolver loss by maintaining traffic that prevents 100% window loss. |
| [`cmd/client/main.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/cmd/client/main.go) | CLI entry point showing balancer initialization and flag binding for health-check parameters. |
| `README.MD` | Documents the `--auto-disable` flag and configuration file parameters. |

## Summary

- MasterDnsVPN tracks every DNS query as a health probe in `balancerResolverSample` structures within [`internal/client/balancer.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/balancer.go).
- Auto-disable triggers only when `autoDisableEnabled` is true and the resolver hits 100% loss across `autoDisableMinObservationsForActiveCount` probes.
- Use `SetResolverDownConfirmHandler` to intercept disable decisions and inject custom validation logic.
- Failed resolvers can automatically recover if `RecheckInactiveServersEnabled` is true and the MTU probe in `recheckInactiveResolver` succeeds.
- Debug by inspecting `connectionStats` counters, monitoring `CollectExpiredResolverTimeouts`, and temporarily shortening `resolverHealthProbeTimeout` to force rapid failures.

## Frequently Asked Questions

### Why does my resolver disable immediately after a single timeout?

When fewer than three resolvers are active, `autoDisableMinObservationsForActiveCount` returns a value of 1, meaning the balancer disables the resolver after the first lost probe. Increase the number of healthy resolvers or implement a custom `SetResolverDownConfirmHandler` to require additional verification before disablement.

### How can I prevent the balancer from ever disabling resolvers?

Call `SetAutoDisableConfig(false, 0)` during initialization or omit the `--auto-disable` CLI flag. When disabled, the balancer continues to track loss statistics in `connectionStats` but never invokes `SetConnectionValidityWithLog` to mark resolvers invalid.

### Why aren’t my disabled resolvers rejoining the active pool?

Verify that `cfg.RecheckInactiveServersEnabled` is set to `true` in your configuration. The background loop in [`internal/client/mtu.go`](https://github.com/masterking32/MasterDnsVPN/blob/main/internal/client/mtu.go) must be running to execute `recheckInactiveResolver`. Also ensure `resolverHealthRecheckInterval` is not set to an excessively long duration that delays re-testing.

### What is the difference between `TrackResolverFailure` and `ReportTimeout`?

`TrackResolverFailure` is the entry point called when a specific DNS response times out; it manages the `balancerResolverSample` lifecycle. `ReportTimeout` is the aggregation layer that increments loss counters and evaluates whether the sliding-window threshold for auto-disable has been reached.