How to Debug MasterDnsVPN Resolver Health Check Failures and Auto-Disable Behavior
MasterDnsVPN disables DNS resolvers when they reach a 100% loss rate inside a configurable sliding window, but you can debug unexpected dropouts by inspecting connectionStats in internal/client/balancer.go and verifying the autoDisableMinObservationsForActiveCount threshold.
MasterDnsVPN continuously probes its upstream DNS resolvers to ensure they remain reachable and performant. When a resolver repeatedly fails to answer a probe, the balancer marks it as invalid and optionally auto-disables it to prevent traffic from routing to a dead endpoint. Understanding the three-part health-check architecture—probe scheduling, timeout aggregation, and background re-checking—is essential for diagnosing why resolvers disappear from the active set.
How the Health-Check System Works
The health-check system is split across three distinct components that track probe lifecycle from transmission to auto-disable decision.
Probe Scheduling and Tracking
Every outbound DNS query triggers TrackResolverSend in internal/client/balancer.go. This method stores a balancerResolverSample keyed by resolver address, local address, and DNS-ID, recording the exact send time. If a response arrives, TrackResolverSuccess removes the sample and retracts any previously recorded timeout via RetractTimeout. If no response arrives before the deadline, TrackResolverFailure invokes ReportTimeout, which increments the lost counter in the resolver’s connectionStats structure.
Timeout Aggregation and Auto-Disable Logic
ReportTimeout aggregates failures into a sliding-window statistic. It checks whether the resolver has sent enough probes to meet the minimum observation count (autoDisableMinObservationsForActiveCount) and whether all probes within the current window have timed out. When both conditions are met, the code may invoke an optional confirmResolverDown handler. If that handler returns true (or is nil), the balancer marks the resolver invalid via SetConnectionValidityWithLog and triggers onResolverDisabled.
Background Health Checks for Inactive Resolvers
When cfg.RecheckInactiveServersEnabled is true, runResolverHealthLoop in internal/client/mtu.go periodically selects inactive resolvers using NextInactiveConnectionForHealthCheck. It runs a full MTU probe (recheckInactiveResolver) against each candidate. A successful MTU test re-enables the resolver by calling SetConnectionValidityWithLog(true, true), while failures keep it disabled until the next iteration.
Common Reasons for Unexpected Auto-Disable
Several configuration and environmental factors can cause premature or confusing resolver disablement:
-
autoDisableEnabledis false – When disabled, the balancer records loss statistics but never callsSetConnectionValidityWithLogto disable the resolver. Verify the flag viabalancer.autoDisableEnabledor by checking the startup logs for theSetAutoDisableConfiginitialization. -
Low active resolver count – The function
autoDisableMinObservationsForActiveCount(active, window)returns a very low threshold (often 1) when fewer than 3 resolvers are active. This causes immediate disablement after a single timeout rather than tolerating transient packet loss. -
Missing confirmation handler – If
confirmResolverDownis nil, the balancer disables the resolver immediately upon hitting the 100% loss threshold without secondary verification. Set a custom handler withSetResolverDownConfirmHandlerto inject an extra probe or logging step. -
Disabled re-check loop – Inactive resolvers never re-enable if
RecheckInactiveServersEnabledis false or ifresolverHealthRecheckIntervalis set too conservatively. Ensure the background health loop is running and the poll interval (resolverHealthPollInterval) is sufficiently aggressive.
Step-by-Step Debugging Procedures
Follow these steps to isolate the exact failure point in the health-check pipeline:
-
Enable debug-level logging
Set the balancer logger toDEBUGto capture per-probe outcomes:balancer.log.SetLevel(logger.LevelDebug) // Or via CLI: --log-level=debug -
Inspect sliding-window statistics
Dump the raw counters for a specific resolver key to see sent vs. lost ratios:stats := balancer.statsForKey("resolver-key-here") fmt.Printf("sent=%d lost=%d windowSent=%d windowLost=%d\n", stats.sent.Load(), stats.lost.Load(), stats.windowSent.Load(), stats.windowLost.Load()) -
Force a rapid timeout
Temporarily reduceresolverHealthProbeTimeoutto 200 ms and enable auto-disable with a short window to reproduce the failure quickly:balancer.SetAutoDisableConfig(true, 2*time.Second) -
Monitor pending-sample sweeps
Add instrumentation at the start ofCollectExpiredResolverTimeoutsto see how many stale probes accumulate. High pending counts indicate the success path is failing to clean up entries. -
Implement a confirmation logger
Inject a confirmation handler to audit the disable decision:balancer.SetResolverDownConfirmHandler(func(c *client.Connection, w time.Duration) bool { fmt.Printf("confirm down for %s (window %v)\n", c.Key, w) return true // return false to block disablement }) -
Trigger a manual re-check
Test an inactive resolver directly to verify the re-enable path:conn, _ := balancer.GetConnectionByKey("resolver-key") client.recheckInactiveResolver(context.Background(), conn)Look for log lines containing
✅ Accepted(MTU success) followed by🟢balancer messages indicating the resolver returned to the active pool.
Practical Example: Diagnosing a Disabled Resolver
Use this debug function immediately after observing a "DNS Resolver disabled" warning to determine whether the resolver met the all-timeouts rule or was forced by a custom handler:
func debugResolver(key string, b *client.Balancer) {
// 1. Verify auto-disable is active
fmt.Printf("Auto-disable enabled: %v\n", b.autoDisableEnabled)
// 2. Calculate the observation window for current active count
active := b.ActiveCount()
window := time.Duration(b.autoDisableTimeoutWindow) * time.Second
fmt.Printf("Active resolvers: %d, window: %v\n", active, window)
// 3. Dump loss statistics
if stats := b.statsForKey(key); stats != nil {
fmt.Printf("sent=%d lost=%d\n", stats.sent.Load(), stats.lost.Load())
}
// 4. Check pending sample queue depth
fmt.Printf("Pending samples: %d\n", b.pendingCount())
}
Call debugResolver("<resolver-key>", balancer) to see if the lost count equals the sent count within the calculated window, confirming the 100% loss trigger fired.
Critical Configuration Parameters
These methods and flags control sensitivity and recovery behavior:
-
SetAutoDisableConfig(enabled bool, window time.Duration)– Enables auto-disable and sets the sliding-window length (typically5sto30s) over which loss is measured. -
SetResolverDisabledHandler(fn func(*Connection, string))– Invoked when a resolver is marked invalid. Use this to emit metrics or pager alerts. -
SetResolverDownConfirmHandler(fn func(*Connection, time.Duration) bool)– Optional gatekeeper that executes before disablement. Returntrueto allow the disable,falseto keep the resolver active. -
cfg.RecheckInactiveServersEnabled– Boolean toggle for the background health loop ininternal/client/mtu.gothat attempts to revive failed resolvers. -
resolverHealthRecheckInterval/resolverHealthPollInterval– Control how frequently inactive resolvers are re-probed and how long the health loop sleeps between sweeps.
Key Source Files
| File | Purpose |
|---|---|
internal/client/balancer.go |
Core balancing, probe tracking via TrackResolverSend, timeout aggregation in ReportTimeout, and auto-disable logic. |
internal/client/mtu.go |
Background health loop (runResolverHealthLoop), inactive resolver selection (NextInactiveConnectionForHealthCheck), and re-enablement via MTU probes. |
internal/client/ping_manager.go |
Controls ping intervals; high-frequency pings can mask resolver loss by maintaining traffic that prevents 100% window loss. |
cmd/client/main.go |
CLI entry point showing balancer initialization and flag binding for health-check parameters. |
README.MD |
Documents the --auto-disable flag and configuration file parameters. |
Summary
- MasterDnsVPN tracks every DNS query as a health probe in
balancerResolverSamplestructures withininternal/client/balancer.go. - Auto-disable triggers only when
autoDisableEnabledis true and the resolver hits 100% loss acrossautoDisableMinObservationsForActiveCountprobes. - Use
SetResolverDownConfirmHandlerto intercept disable decisions and inject custom validation logic. - Failed resolvers can automatically recover if
RecheckInactiveServersEnabledis true and the MTU probe inrecheckInactiveResolversucceeds. - Debug by inspecting
connectionStatscounters, monitoringCollectExpiredResolverTimeouts, and temporarily shorteningresolverHealthProbeTimeoutto force rapid failures.
Frequently Asked Questions
Why does my resolver disable immediately after a single timeout?
When fewer than three resolvers are active, autoDisableMinObservationsForActiveCount returns a value of 1, meaning the balancer disables the resolver after the first lost probe. Increase the number of healthy resolvers or implement a custom SetResolverDownConfirmHandler to require additional verification before disablement.
How can I prevent the balancer from ever disabling resolvers?
Call SetAutoDisableConfig(false, 0) during initialization or omit the --auto-disable CLI flag. When disabled, the balancer continues to track loss statistics in connectionStats but never invokes SetConnectionValidityWithLog to mark resolvers invalid.
Why aren’t my disabled resolvers rejoining the active pool?
Verify that cfg.RecheckInactiveServersEnabled is set to true in your configuration. The background loop in internal/client/mtu.go must be running to execute recheckInactiveResolver. Also ensure resolverHealthRecheckInterval is not set to an excessively long duration that delays re-testing.
What is the difference between TrackResolverFailure and ReportTimeout?
TrackResolverFailure is the entry point called when a specific DNS response times out; it manages the balancerResolverSample lifecycle. ReportTimeout is the aggregation layer that increments loss counters and evaluates whether the sliding-window threshold for auto-disable has been reached.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →