How to Debug MasterDnsVPN Resolver Health Check Failures and Auto-Disable Behavior

MasterDnsVPN disables DNS resolvers when they reach a 100% loss rate inside a configurable sliding window, but you can debug unexpected dropouts by inspecting connectionStats in internal/client/balancer.go and verifying the autoDisableMinObservationsForActiveCount threshold.

MasterDnsVPN continuously probes its upstream DNS resolvers to ensure they remain reachable and performant. When a resolver repeatedly fails to answer a probe, the balancer marks it as invalid and optionally auto-disables it to prevent traffic from routing to a dead endpoint. Understanding the three-part health-check architecture—probe scheduling, timeout aggregation, and background re-checking—is essential for diagnosing why resolvers disappear from the active set.

How the Health-Check System Works

The health-check system is split across three distinct components that track probe lifecycle from transmission to auto-disable decision.

Probe Scheduling and Tracking

Every outbound DNS query triggers TrackResolverSend in internal/client/balancer.go. This method stores a balancerResolverSample keyed by resolver address, local address, and DNS-ID, recording the exact send time. If a response arrives, TrackResolverSuccess removes the sample and retracts any previously recorded timeout via RetractTimeout. If no response arrives before the deadline, TrackResolverFailure invokes ReportTimeout, which increments the lost counter in the resolver’s connectionStats structure.

Timeout Aggregation and Auto-Disable Logic

ReportTimeout aggregates failures into a sliding-window statistic. It checks whether the resolver has sent enough probes to meet the minimum observation count (autoDisableMinObservationsForActiveCount) and whether all probes within the current window have timed out. When both conditions are met, the code may invoke an optional confirmResolverDown handler. If that handler returns true (or is nil), the balancer marks the resolver invalid via SetConnectionValidityWithLog and triggers onResolverDisabled.

Background Health Checks for Inactive Resolvers

When cfg.RecheckInactiveServersEnabled is true, runResolverHealthLoop in internal/client/mtu.go periodically selects inactive resolvers using NextInactiveConnectionForHealthCheck. It runs a full MTU probe (recheckInactiveResolver) against each candidate. A successful MTU test re-enables the resolver by calling SetConnectionValidityWithLog(true, true), while failures keep it disabled until the next iteration.

Common Reasons for Unexpected Auto-Disable

Several configuration and environmental factors can cause premature or confusing resolver disablement:

  • autoDisableEnabled is false – When disabled, the balancer records loss statistics but never calls SetConnectionValidityWithLog to disable the resolver. Verify the flag via balancer.autoDisableEnabled or by checking the startup logs for the SetAutoDisableConfig initialization.

  • Low active resolver count – The function autoDisableMinObservationsForActiveCount(active, window) returns a very low threshold (often 1) when fewer than 3 resolvers are active. This causes immediate disablement after a single timeout rather than tolerating transient packet loss.

  • Missing confirmation handler – If confirmResolverDown is nil, the balancer disables the resolver immediately upon hitting the 100% loss threshold without secondary verification. Set a custom handler with SetResolverDownConfirmHandler to inject an extra probe or logging step.

  • Disabled re-check loop – Inactive resolvers never re-enable if RecheckInactiveServersEnabled is false or if resolverHealthRecheckInterval is set too conservatively. Ensure the background health loop is running and the poll interval (resolverHealthPollInterval) is sufficiently aggressive.

Step-by-Step Debugging Procedures

Follow these steps to isolate the exact failure point in the health-check pipeline:

  1. Enable debug-level logging
    Set the balancer logger to DEBUG to capture per-probe outcomes:

    balancer.log.SetLevel(logger.LevelDebug)
    // Or via CLI: --log-level=debug
  2. Inspect sliding-window statistics
    Dump the raw counters for a specific resolver key to see sent vs. lost ratios:

    stats := balancer.statsForKey("resolver-key-here")
    fmt.Printf("sent=%d lost=%d windowSent=%d windowLost=%d\n",
        stats.sent.Load(), stats.lost.Load(),
        stats.windowSent.Load(), stats.windowLost.Load())
  3. Force a rapid timeout
    Temporarily reduce resolverHealthProbeTimeout to 200 ms and enable auto-disable with a short window to reproduce the failure quickly:

    balancer.SetAutoDisableConfig(true, 2*time.Second)
  4. Monitor pending-sample sweeps
    Add instrumentation at the start of CollectExpiredResolverTimeouts to see how many stale probes accumulate. High pending counts indicate the success path is failing to clean up entries.

  5. Implement a confirmation logger
    Inject a confirmation handler to audit the disable decision:

    balancer.SetResolverDownConfirmHandler(func(c *client.Connection, w time.Duration) bool {
        fmt.Printf("confirm down for %s (window %v)\n", c.Key, w)
        return true // return false to block disablement
    })
  6. Trigger a manual re-check
    Test an inactive resolver directly to verify the re-enable path:

    conn, _ := balancer.GetConnectionByKey("resolver-key")
    client.recheckInactiveResolver(context.Background(), conn)

    Look for log lines containing ✅ Accepted (MTU success) followed by 🟢 balancer messages indicating the resolver returned to the active pool.

Practical Example: Diagnosing a Disabled Resolver

Use this debug function immediately after observing a "DNS Resolver disabled" warning to determine whether the resolver met the all-timeouts rule or was forced by a custom handler:

func debugResolver(key string, b *client.Balancer) {
    // 1. Verify auto-disable is active
    fmt.Printf("Auto-disable enabled: %v\n", b.autoDisableEnabled)

    // 2. Calculate the observation window for current active count
    active := b.ActiveCount()
    window := time.Duration(b.autoDisableTimeoutWindow) * time.Second
    fmt.Printf("Active resolvers: %d, window: %v\n", active, window)

    // 3. Dump loss statistics
    if stats := b.statsForKey(key); stats != nil {
        fmt.Printf("sent=%d lost=%d\n", stats.sent.Load(), stats.lost.Load())
    }

    // 4. Check pending sample queue depth
    fmt.Printf("Pending samples: %d\n", b.pendingCount())
}

Call debugResolver("<resolver-key>", balancer) to see if the lost count equals the sent count within the calculated window, confirming the 100% loss trigger fired.

Critical Configuration Parameters

These methods and flags control sensitivity and recovery behavior:

  • SetAutoDisableConfig(enabled bool, window time.Duration) – Enables auto-disable and sets the sliding-window length (typically 5s to 30s) over which loss is measured.

  • SetResolverDisabledHandler(fn func(*Connection, string)) – Invoked when a resolver is marked invalid. Use this to emit metrics or pager alerts.

  • SetResolverDownConfirmHandler(fn func(*Connection, time.Duration) bool) – Optional gatekeeper that executes before disablement. Return true to allow the disable, false to keep the resolver active.

  • cfg.RecheckInactiveServersEnabled – Boolean toggle for the background health loop in internal/client/mtu.go that attempts to revive failed resolvers.

  • resolverHealthRecheckInterval / resolverHealthPollInterval – Control how frequently inactive resolvers are re-probed and how long the health loop sleeps between sweeps.

Key Source Files

File Purpose
internal/client/balancer.go Core balancing, probe tracking via TrackResolverSend, timeout aggregation in ReportTimeout, and auto-disable logic.
internal/client/mtu.go Background health loop (runResolverHealthLoop), inactive resolver selection (NextInactiveConnectionForHealthCheck), and re-enablement via MTU probes.
internal/client/ping_manager.go Controls ping intervals; high-frequency pings can mask resolver loss by maintaining traffic that prevents 100% window loss.
cmd/client/main.go CLI entry point showing balancer initialization and flag binding for health-check parameters.
README.MD Documents the --auto-disable flag and configuration file parameters.

Summary

  • MasterDnsVPN tracks every DNS query as a health probe in balancerResolverSample structures within internal/client/balancer.go.
  • Auto-disable triggers only when autoDisableEnabled is true and the resolver hits 100% loss across autoDisableMinObservationsForActiveCount probes.
  • Use SetResolverDownConfirmHandler to intercept disable decisions and inject custom validation logic.
  • Failed resolvers can automatically recover if RecheckInactiveServersEnabled is true and the MTU probe in recheckInactiveResolver succeeds.
  • Debug by inspecting connectionStats counters, monitoring CollectExpiredResolverTimeouts, and temporarily shortening resolverHealthProbeTimeout to force rapid failures.

Frequently Asked Questions

Why does my resolver disable immediately after a single timeout?

When fewer than three resolvers are active, autoDisableMinObservationsForActiveCount returns a value of 1, meaning the balancer disables the resolver after the first lost probe. Increase the number of healthy resolvers or implement a custom SetResolverDownConfirmHandler to require additional verification before disablement.

How can I prevent the balancer from ever disabling resolvers?

Call SetAutoDisableConfig(false, 0) during initialization or omit the --auto-disable CLI flag. When disabled, the balancer continues to track loss statistics in connectionStats but never invokes SetConnectionValidityWithLog to mark resolvers invalid.

Why aren’t my disabled resolvers rejoining the active pool?

Verify that cfg.RecheckInactiveServersEnabled is set to true in your configuration. The background loop in internal/client/mtu.go must be running to execute recheckInactiveResolver. Also ensure resolverHealthRecheckInterval is not set to an excessively long duration that delays re-testing.

What is the difference between TrackResolverFailure and ReportTimeout?

TrackResolverFailure is the entry point called when a specific DNS response times out; it manages the balancerResolverSample lifecycle. ReportTimeout is the aggregation layer that increments loss counters and evaluates whether the sliding-window threshold for auto-disable has been reached.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →