How Experiment Parameters Are Synchronized From Production Config in X-Algorithm

Experiment parameters are automatically synchronized from production configuration by reading the config.json checkpoint file, deserializing it into a typed object with xai_configlib, and broadcasting the resulting configuration to all distributed training hosts before training begins.

In the xai-org/x-algorithm repository, experiment-time hyper-parameters remain tightly coupled to production-time settings through a deterministic synchronization mechanism. This process ensures that when researchers resume training from production checkpoints, they inherit exactly the same hyper-parameters used in the original service, eliminating configuration drift across distributed environments.

The Three-Step Synchronization Pipeline

The synchronization process implemented in phoenix/xrex/train/trainer_recsys.py consists of three tightly-coupled steps that guarantee configuration consistency across distributed training jobs.

Loading the Production Config from Checkpoint

When a training job starts or resumes from a checkpoint, the trainer locates and reads the config.json file written by the production service. This occurs in the _read_checkpoint_kafka_config helper method inside phoenix/xrex/train/trainer_recsys.py.


# phoenix/xrex/train/trainer_recsys.py

def _read_checkpoint_kafka_config(ctx) -> tuple[str | None, int | None]:
    config_path = os.path.join(ctx.checkpoint.path, _CONFIG_FILENAME)
    if not os.path.isfile(config_path):
        return None, None
    with open(config_path) as f:
        config = json.load(f)

    # The file may contain a nested "dataset" section; pull it out.

    dataset_config = config.get("dataset", config)
    topic = dataset_config.get("topic_name")
    partitions = dataset_config.get("num_kafka_partitions")
    return topic, partitions

The JSON file contains the full xai_configlib configuration hierarchy, including dataset, model, and optimizer settings.

Deserializing with xai_configlib

After reading the raw JSON, the system deserializes it into a typed configuration object using xai_configlib.Config.from_json. This materializes a Python class decorated with @configclass, ensuring that the same structure definitions used in production are enforced during experimentation.


# phoenix/xrex/train/trainer_recsys.py (excerpt)

from xai_configlib import Config
from xrex.utils.utils import broadcast_one_to_all

# After reading config.json …

raw_cfg = json.load(open(config_path))
prod_cfg = Config.from_json(raw_cfg)                # Typed config object

Because the codebase shares configclass definitions between production and experiment environments, the structures are guaranteed to match, preventing type mismatches or missing fields.

Broadcasting to All Hosts

In multi-host training scenarios, the configuration must be identical on every process. After the root process deserializes the configuration, it propagates the object to all participants using broadcast_one_to_all.


# Continued from previous excerpt

prod_cfg = broadcast_one_to_all(prod_cfg)          # Sync across hosts

Under the hood, phoenix/xrex/utils/utils.py wraps jax.experimental.multihost_utils.broadcast_one_to_all to handle the cross-host communication. This operation completes before any training step begins, ensuring every host sees the exact same hyper-parameters.

Persisting Configuration for Reproducibility

The metadata subsystem continuously records the active configuration throughout training, storing a JSON copy in the checkpoint directory. This creates a single source of truth for both production runs and experimental reruns.

In phoenix/xrex/utils/metadata.py, the record_run method writes the configuration on every checkpoint:


# phoenix/xrex/utils/metadata.py

def record_run(self, run: Run, config: Jsonable):
    """Write a copy of the config to the checkpoint directory."""
    with (path / "config.json").open("w") as f:
        f.write(self.current_config.to_json())

This persistence mechanism makes it trivial to resume jobs with identical parameters and audit the exact configuration used for any given checkpoint.

Summary

  • Automatic synchronization ensures experiment parameters match production configs by loading config.json from checkpoints in phoenix/xrex/train/trainer_recsys.py.
  • Type-safe deserialization uses xai_configlib.Config.from_json to materialize typed configuration objects that share schema definitions with production.
  • Distributed consistency is achieved via broadcast_one_to_all in phoenix/xrex/utils/utils.py, which copies the config to all training hosts before execution.
  • Reproducibility is guaranteed by the metadata system in phoenix/xrex/utils/metadata.py, which persists configurations alongside checkpoints.

Frequently Asked Questions

How does X-Algorithm ensure all training hosts use identical parameters?

The system uses multihost_utils.broadcast_one_to_all (wrapped by xrex.utils.utils.broadcast_one_to_all) to propagate the deserialized configuration from the root process to all distributed hosts immediately after loading. This broadcast occurs before training begins, ensuring every participant starts with the exact same hyper-parameters.

What file format is used to store production configuration?

Configurations are stored as JSON files named config.json located within the checkpoint directory. This file contains the complete xai_configlib hierarchy including dataset, model, and optimizer settings, written by the production service when the checkpoint is created.

Where is the configuration synchronization logic implemented?

The primary logic resides in phoenix/xrex/train/trainer_recsys.py, specifically within the _read_checkpoint_kafka_config method for loading and the main training setup code for deserializing and broadcasting. Supporting utilities for multi-host broadcasting live in phoenix/xrex/utils/utils.py.

How does xai_configlib maintain type safety during config synchronization?

The library uses Python classes decorated with @configclass to define configuration schemas. When Config.from_json deserializes the checkpoint's JSON, it materializes instances of these typed classes rather than raw dictionaries. Because production and experiment code share the same configclass definitions, the structure and types are guaranteed to match between environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →