Xray-core Observatory Module: Health Checking and Outbound Monitoring Explained
The Xray-core Observatory module is an internal application that continuously monitors outbound proxy health through HTTP probing and statistical ping analysis, exposing real-time status data to routing balancers for intelligent traffic distribution.
The Xray-core Observatory module provides infrastructure-level visibility into outbound proxy performance within the XTLS/Xray-core platform. As an implementation of the features.Extension interface, it registers protobuf-based configurations and executes automated health checks against configured outbound tags. This system enables dynamic routing decisions by supplying latency statistics and availability metrics to load balancing strategies like strategy_leastload.go.
What Is the Xray-core Observatory Module?
The Observatory operates as a core application within app/observatory/, implementing continuous health monitoring for outbound handlers. It exposes two distinct observation patterns through a unified interface defined in features/extension/observatory.go:
Simple Observer
The Simple Observer performs single HTTP GET requests to validate outbound connectivity. Implemented in app/observatory/observer.go, this component probes each configured outbound tag at regular intervals and returns immediate OutboundStatus results through the GetObservation method.
Burst Observer
The Burst Observer provides statistical health analysis through repeated ping sampling. Located in app/observatory/burst/burstobserver.go, this implementation aggregates round-trip time (RTT) data, calculates deviations, and tracks success rates over configurable windows, offering more sophisticated health metrics than simple binary up/down checks.
How Health Checking Works in Xray-core
The module executes health checks through two distinct mechanisms, both leveraging the tagged.Dialer system to force traffic through specific outbound tags.
Simple HTTP Probing Implementation
The simple probing mechanism follows a straightforward execution flow defined in app/observatory/observer.go:
-
Configuration Loading: The
ObservatoryConfigstruct (infra/conf/observatory.go) parsessubjectSelector,probeURL(defaulting tohttps://www.google.com/generate_204),probeInterval, and concurrency settings. -
Background Execution: The
Observer.backgroundmethod initiates a ticker loop that selects outbound tags viaoutbound.HandlerSelectorusinghs.Select(o.config.SubjectSelector)at line 72. -
Tagged Transport Creation: For each selected tag,
Observer.probeconstructs a custom HTTP transport that routes through the specific outbound usingtagged.Dialer(lines 46-48). -
Latency Recording: Successful requests update an internal
OutboundStatusmap with measured delays, accessible viaGetObservation.
Burst Health-Ping Mechanism
The Burst Observer implements statistical health analysis through the HealthPing subsystem in app/observatory/burst/healthping.go:
-
Scheduler Initialization:
BurstObserver.Startinvokeshp.StartScheduler(lines 70-82 inburstobserver.go), creating a ticker based onInterval * SamplingCount. -
Concurrent Ping Execution: The
HealthPing.doCheckmethod (lines 61-84) launches parallel ping tasks for each configured tag. Each execution uses apingClientto establish connections through the selected outbound and execute HTTP requests (defaulting to HEAD method). -
Statistical Aggregation: Results feed into
HealthPingRTTSbuffers (healthping_result.go), which maintain sliding windows of RTT samples. TheHealthPingRTTS.Getmethod returnsHealthPingStatscontainingAll,Fail,Average,Deviation,Max, andMinvalues. -
Result Compilation:
BurstObserver.createResultwalks thehp.Resultsmap to generateobservatory.OutboundStatusslices enriched withHealthPingMeasurementResultprotobuf messages.
Integration with Load Balancers
Health metrics flow into routing decisions through app/router/strategy_leastload.go. The least-load strategy reads OutboundStatus.HealthPing fields to calculate routing weights:
if v.HealthPing != nil {
record.RTTAverage = time.Duration(v.HealthPing.Average)
record.RTTDeviation = time.Duration(v.HealthPing.Deviation)
}
This integration allows balancers to prefer outbounds with lower average latency and stable deviation patterns.
Configuration and Implementation Examples
Configuring Simple HTTP Observation
Configure basic health monitoring in your Xray JSON configuration:
{
"observatory": {
"subjectSelector": ["proxy-us", "proxy-sg"],
"probeURL": "https://connectivitycheck.gstatic.com/generate_204",
"probeInterval": "30s",
"enableConcurrency": true
}
}
This maps to ObservatoryConfig in infra/conf/observatory.go, creating probes that execute every 30 seconds against the specified outbounds.
Configuring Burst Health-Ping Monitoring
For statistical analysis with sampling, use the burst configuration:
{
"burstObservatory": {
"subjectSelector": ["proxy-us", "proxy-sg"],
"pingConfig": {
"destination": "https://connectivitycheck.gstatic.com/generate_204",
"interval": "60s",
"samplingCount": 10,
"timeout": "5s",
"httpMethod": "HEAD"
}
}
}
Parsed by BurstObservatoryConfig.Build() (lines 29-38 in infra/conf/observatory.go), this configuration triggers the statistical aggregation pipeline in the burst observer.
Querying Health Data Programmatically
Access observation results from Go code using the feature system:
import (
"context"
"fmt"
"github.com/xtls/xray-core/app/observatory"
"github.com/xtls/xray-core/core"
"github.com/xtls/xray-core/features/extension"
)
func checkOutboundHealth() {
ctx := core.NewContext(context.Background())
obs, err := core.RequireFeature(ctx, func(o extension.Observatory) {})
if err != nil {
panic(err)
}
msg, err := obs.GetObservation(ctx)
if err != nil {
panic(err)
}
result := msg.(*observatory.ObservationResult)
for _, s := range result.Status {
fmt.Printf("Tag=%s Alive=%v Delay=%dms HealthPing-Avg=%d\n",
s.OutboundTag, s.Alive, s.Delay, s.HealthPing.Average)
}
}
Custom Balancer Implementation
Consume health-ping data in custom routing logic:
func selectLowestLatency(outbounds []string, obs *observatory.ObservationResult) string {
var best string
var bestRTT int64 = 1<<63 - 1
for _, st := range obs.Status {
if st.Alive && st.HealthPing != nil && st.HealthPing.Average < bestRTT {
bestRTT = st.HealthPing.Average
best = st.OutboundTag
}
}
return best
}
Summary
- The Xray-core Observatory module implements the
features.Extensioninterface to provide continuous outbound health monitoring throughapp/observatory/observer.goandapp/observatory/burst/burstobserver.go. - Simple Observer executes single HTTP probes via
tagged.Dialerto verify basic connectivity and measure latency. - Burst Observer runs statistical health-ping checks through
HealthPing.doCheck, aggregating RTT samples into deviation and average metrics stored inHealthPingRTTS. - Configuration occurs through
infra/conf/observatory.go, supporting both immediate probing and statistical sampling via protobuf-defined settings. - Health data integrates with routing strategies like
strategy_leastload.go, enabling latency-aware traffic distribution based on real-timeHealthPingMeasurementResultdata.
Frequently Asked Questions
What is the difference between Simple and Burst observers in Xray-core?
The Simple Observer performs single HTTP GET requests at configured intervals to determine basic availability and latency, storing results in OutboundStatus. The Burst Observer collects multiple samples over time windows, calculating statistical metrics including average RTT, deviation, and failure rates through HealthPingStats, providing more reliable health assessments for load-balancing decisions.
How does the Observatory module integrate with load balancers?
The module exposes health data through the GetObservation interface, which returns ObservationResult containing HealthPing fields. Balancers like the least-load strategy in app/router/strategy_leastload.go read HealthPing.Average and HealthPing.Deviation to weight routing decisions, preferring outbounds with lower latency and stable performance characteristics.
What URL should I use for the probeURL destination?
While the default configuration uses https://www.google.com/generate_204, production deployments should specify reliable endpoints like https://connectivitycheck.gstatic.com/generate_204 or other infrastructure-specific health check URLs. The destination must return HTTP 204 or success codes quickly to ensure accurate latency measurements without processing overhead.
Can I query Observatory health data programmatically?
Yes. The module registers as a feature accessible through core.RequireFeature with the extension.Observatory interface. Applications can call GetObservation(context.Context) to receive ObservationResult protobuf messages containing current OutboundStatus for all monitored tags, including latency data and health-ping statistics when using the burst observer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →