How Experiment Parameters Are Synchronized From Production Config in X-Algorithm
Experiment parameters are automatically synchronized from production configuration by reading the config.json checkpoint file, deserializing it into a typed object with xai_configlib, and broadcasting the resulting configuration to all distributed training hosts before training begins.
In the xai-org/x-algorithm repository, experiment-time hyper-parameters remain tightly coupled to production-time settings through a deterministic synchronization mechanism. This process ensures that when researchers resume training from production checkpoints, they inherit exactly the same hyper-parameters used in the original service, eliminating configuration drift across distributed environments.
The Three-Step Synchronization Pipeline
The synchronization process implemented in phoenix/xrex/train/trainer_recsys.py consists of three tightly-coupled steps that guarantee configuration consistency across distributed training jobs.
Loading the Production Config from Checkpoint
When a training job starts or resumes from a checkpoint, the trainer locates and reads the config.json file written by the production service. This occurs in the _read_checkpoint_kafka_config helper method inside phoenix/xrex/train/trainer_recsys.py.
# phoenix/xrex/train/trainer_recsys.py
def _read_checkpoint_kafka_config(ctx) -> tuple[str | None, int | None]:
config_path = os.path.join(ctx.checkpoint.path, _CONFIG_FILENAME)
if not os.path.isfile(config_path):
return None, None
with open(config_path) as f:
config = json.load(f)
# The file may contain a nested "dataset" section; pull it out.
dataset_config = config.get("dataset", config)
topic = dataset_config.get("topic_name")
partitions = dataset_config.get("num_kafka_partitions")
return topic, partitions
The JSON file contains the full xai_configlib configuration hierarchy, including dataset, model, and optimizer settings.
Deserializing with xai_configlib
After reading the raw JSON, the system deserializes it into a typed configuration object using xai_configlib.Config.from_json. This materializes a Python class decorated with @configclass, ensuring that the same structure definitions used in production are enforced during experimentation.
# phoenix/xrex/train/trainer_recsys.py (excerpt)
from xai_configlib import Config
from xrex.utils.utils import broadcast_one_to_all
# After reading config.json …
raw_cfg = json.load(open(config_path))
prod_cfg = Config.from_json(raw_cfg) # Typed config object
Because the codebase shares configclass definitions between production and experiment environments, the structures are guaranteed to match, preventing type mismatches or missing fields.
Broadcasting to All Hosts
In multi-host training scenarios, the configuration must be identical on every process. After the root process deserializes the configuration, it propagates the object to all participants using broadcast_one_to_all.
# Continued from previous excerpt
prod_cfg = broadcast_one_to_all(prod_cfg) # Sync across hosts
Under the hood, phoenix/xrex/utils/utils.py wraps jax.experimental.multihost_utils.broadcast_one_to_all to handle the cross-host communication. This operation completes before any training step begins, ensuring every host sees the exact same hyper-parameters.
Persisting Configuration for Reproducibility
The metadata subsystem continuously records the active configuration throughout training, storing a JSON copy in the checkpoint directory. This creates a single source of truth for both production runs and experimental reruns.
In phoenix/xrex/utils/metadata.py, the record_run method writes the configuration on every checkpoint:
# phoenix/xrex/utils/metadata.py
def record_run(self, run: Run, config: Jsonable):
"""Write a copy of the config to the checkpoint directory."""
with (path / "config.json").open("w") as f:
f.write(self.current_config.to_json())
This persistence mechanism makes it trivial to resume jobs with identical parameters and audit the exact configuration used for any given checkpoint.
Summary
- Automatic synchronization ensures experiment parameters match production configs by loading
config.jsonfrom checkpoints inphoenix/xrex/train/trainer_recsys.py. - Type-safe deserialization uses
xai_configlib.Config.from_jsonto materialize typed configuration objects that share schema definitions with production. - Distributed consistency is achieved via
broadcast_one_to_allinphoenix/xrex/utils/utils.py, which copies the config to all training hosts before execution. - Reproducibility is guaranteed by the metadata system in
phoenix/xrex/utils/metadata.py, which persists configurations alongside checkpoints.
Frequently Asked Questions
How does X-Algorithm ensure all training hosts use identical parameters?
The system uses multihost_utils.broadcast_one_to_all (wrapped by xrex.utils.utils.broadcast_one_to_all) to propagate the deserialized configuration from the root process to all distributed hosts immediately after loading. This broadcast occurs before training begins, ensuring every participant starts with the exact same hyper-parameters.
What file format is used to store production configuration?
Configurations are stored as JSON files named config.json located within the checkpoint directory. This file contains the complete xai_configlib hierarchy including dataset, model, and optimizer settings, written by the production service when the checkpoint is created.
Where is the configuration synchronization logic implemented?
The primary logic resides in phoenix/xrex/train/trainer_recsys.py, specifically within the _read_checkpoint_kafka_config method for loading and the main training setup code for deserializing and broadcasting. Supporting utilities for multi-host broadcasting live in phoenix/xrex/utils/utils.py.
How does xai_configlib maintain type safety during config synchronization?
The library uses Python classes decorated with @configclass to define configuration schemas. When Config.from_json deserializes the checkpoint's JSON, it materializes instances of these typed classes rather than raw dictionaries. Because production and experiment code share the same configclass definitions, the structure and types are guaranteed to match between environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →