Xray-core Observatory Module: Health Checking and Outbound Monitoring Explained

The Xray-core Observatory module is an internal application that continuously monitors outbound proxy health through HTTP probing and statistical ping analysis, exposing real-time status data to routing balancers for intelligent traffic distribution.

The Xray-core Observatory module provides infrastructure-level visibility into outbound proxy performance within the XTLS/Xray-core platform. As an implementation of the features.Extension interface, it registers protobuf-based configurations and executes automated health checks against configured outbound tags. This system enables dynamic routing decisions by supplying latency statistics and availability metrics to load balancing strategies like strategy_leastload.go.

What Is the Xray-core Observatory Module?

The Observatory operates as a core application within app/observatory/, implementing continuous health monitoring for outbound handlers. It exposes two distinct observation patterns through a unified interface defined in features/extension/observatory.go:

Simple Observer

The Simple Observer performs single HTTP GET requests to validate outbound connectivity. Implemented in app/observatory/observer.go, this component probes each configured outbound tag at regular intervals and returns immediate OutboundStatus results through the GetObservation method.

Burst Observer

The Burst Observer provides statistical health analysis through repeated ping sampling. Located in app/observatory/burst/burstobserver.go, this implementation aggregates round-trip time (RTT) data, calculates deviations, and tracks success rates over configurable windows, offering more sophisticated health metrics than simple binary up/down checks.

How Health Checking Works in Xray-core

The module executes health checks through two distinct mechanisms, both leveraging the tagged.Dialer system to force traffic through specific outbound tags.

Simple HTTP Probing Implementation

The simple probing mechanism follows a straightforward execution flow defined in app/observatory/observer.go:

  1. Configuration Loading: The ObservatoryConfig struct (infra/conf/observatory.go) parses subjectSelector, probeURL (defaulting to https://www.google.com/generate_204), probeInterval, and concurrency settings.

  2. Background Execution: The Observer.background method initiates a ticker loop that selects outbound tags via outbound.HandlerSelector using hs.Select(o.config.SubjectSelector) at line 72.

  3. Tagged Transport Creation: For each selected tag, Observer.probe constructs a custom HTTP transport that routes through the specific outbound using tagged.Dialer (lines 46-48).

  4. Latency Recording: Successful requests update an internal OutboundStatus map with measured delays, accessible via GetObservation.

Burst Health-Ping Mechanism

The Burst Observer implements statistical health analysis through the HealthPing subsystem in app/observatory/burst/healthping.go:

  1. Scheduler Initialization: BurstObserver.Start invokes hp.StartScheduler (lines 70-82 in burstobserver.go), creating a ticker based on Interval * SamplingCount.

  2. Concurrent Ping Execution: The HealthPing.doCheck method (lines 61-84) launches parallel ping tasks for each configured tag. Each execution uses a pingClient to establish connections through the selected outbound and execute HTTP requests (defaulting to HEAD method).

  3. Statistical Aggregation: Results feed into HealthPingRTTS buffers (healthping_result.go), which maintain sliding windows of RTT samples. The HealthPingRTTS.Get method returns HealthPingStats containing All, Fail, Average, Deviation, Max, and Min values.

  4. Result Compilation: BurstObserver.createResult walks the hp.Results map to generate observatory.OutboundStatus slices enriched with HealthPingMeasurementResult protobuf messages.

Integration with Load Balancers

Health metrics flow into routing decisions through app/router/strategy_leastload.go. The least-load strategy reads OutboundStatus.HealthPing fields to calculate routing weights:

if v.HealthPing != nil {
    record.RTTAverage = time.Duration(v.HealthPing.Average)
    record.RTTDeviation = time.Duration(v.HealthPing.Deviation)
}

This integration allows balancers to prefer outbounds with lower average latency and stable deviation patterns.

Configuration and Implementation Examples

Configuring Simple HTTP Observation

Configure basic health monitoring in your Xray JSON configuration:

{
  "observatory": {
    "subjectSelector": ["proxy-us", "proxy-sg"],
    "probeURL": "https://connectivitycheck.gstatic.com/generate_204",
    "probeInterval": "30s",
    "enableConcurrency": true
  }
}

This maps to ObservatoryConfig in infra/conf/observatory.go, creating probes that execute every 30 seconds against the specified outbounds.

Configuring Burst Health-Ping Monitoring

For statistical analysis with sampling, use the burst configuration:

{
  "burstObservatory": {
    "subjectSelector": ["proxy-us", "proxy-sg"],
    "pingConfig": {
      "destination": "https://connectivitycheck.gstatic.com/generate_204",
      "interval": "60s",
      "samplingCount": 10,
      "timeout": "5s",
      "httpMethod": "HEAD"
    }
  }
}

Parsed by BurstObservatoryConfig.Build() (lines 29-38 in infra/conf/observatory.go), this configuration triggers the statistical aggregation pipeline in the burst observer.

Querying Health Data Programmatically

Access observation results from Go code using the feature system:

import (
    "context"
    "fmt"
    "github.com/xtls/xray-core/app/observatory"
    "github.com/xtls/xray-core/core"
    "github.com/xtls/xray-core/features/extension"
)

func checkOutboundHealth() {
    ctx := core.NewContext(context.Background())
    
    obs, err := core.RequireFeature(ctx, func(o extension.Observatory) {})
    if err != nil {
        panic(err)
    }

    msg, err := obs.GetObservation(ctx)
    if err != nil {
        panic(err)
    }

    result := msg.(*observatory.ObservationResult)
    for _, s := range result.Status {
        fmt.Printf("Tag=%s Alive=%v Delay=%dms HealthPing-Avg=%d\n",
            s.OutboundTag, s.Alive, s.Delay, s.HealthPing.Average)
    }
}

Custom Balancer Implementation

Consume health-ping data in custom routing logic:

func selectLowestLatency(outbounds []string, obs *observatory.ObservationResult) string {
    var best string
    var bestRTT int64 = 1<<63 - 1
    
    for _, st := range obs.Status {
        if st.Alive && st.HealthPing != nil && st.HealthPing.Average < bestRTT {
            bestRTT = st.HealthPing.Average
            best = st.OutboundTag
        }
    }
    return best
}

Summary

  • The Xray-core Observatory module implements the features.Extension interface to provide continuous outbound health monitoring through app/observatory/observer.go and app/observatory/burst/burstobserver.go.
  • Simple Observer executes single HTTP probes via tagged.Dialer to verify basic connectivity and measure latency.
  • Burst Observer runs statistical health-ping checks through HealthPing.doCheck, aggregating RTT samples into deviation and average metrics stored in HealthPingRTTS.
  • Configuration occurs through infra/conf/observatory.go, supporting both immediate probing and statistical sampling via protobuf-defined settings.
  • Health data integrates with routing strategies like strategy_leastload.go, enabling latency-aware traffic distribution based on real-time HealthPingMeasurementResult data.

Frequently Asked Questions

What is the difference between Simple and Burst observers in Xray-core?

The Simple Observer performs single HTTP GET requests at configured intervals to determine basic availability and latency, storing results in OutboundStatus. The Burst Observer collects multiple samples over time windows, calculating statistical metrics including average RTT, deviation, and failure rates through HealthPingStats, providing more reliable health assessments for load-balancing decisions.

How does the Observatory module integrate with load balancers?

The module exposes health data through the GetObservation interface, which returns ObservationResult containing HealthPing fields. Balancers like the least-load strategy in app/router/strategy_leastload.go read HealthPing.Average and HealthPing.Deviation to weight routing decisions, preferring outbounds with lower latency and stable performance characteristics.

What URL should I use for the probeURL destination?

While the default configuration uses https://www.google.com/generate_204, production deployments should specify reliable endpoints like https://connectivitycheck.gstatic.com/generate_204 or other infrastructure-specific health check URLs. The destination must return HTTP 204 or success codes quickly to ensure accurate latency measurements without processing overhead.

Can I query Observatory health data programmatically?

Yes. The module registers as a feature accessible through core.RequireFeature with the extension.Observatory interface. Applications can call GetObservation(context.Context) to receive ObservationResult protobuf messages containing current OutboundStatus for all monitored tags, including latency data and health-ping statistics when using the burst observer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →