# How Experiment Parameters Are Synchronized From Production Config in X-Algorithm

> Learn how X-Algorithm synchronizes experiment parameters from production config by reading, deserializing, and broadcasting the config.json checkpoint for distributed training.

- Repository: [SpaceXAI Org/x-algorithm](https://github.com/xai-org/x-algorithm)
- Tags: internals
- Published: 2026-09-11

---

**Experiment parameters are automatically synchronized from production configuration by reading the [`config.json`](https://github.com/xai-org/x-algorithm/blob/main/config.json) checkpoint file, deserializing it into a typed object with `xai_configlib`, and broadcasting the resulting configuration to all distributed training hosts before training begins.**

In the `xai-org/x-algorithm` repository, experiment-time hyper-parameters remain tightly coupled to production-time settings through a deterministic synchronization mechanism. This process ensures that when researchers resume training from production checkpoints, they inherit **exactly the same hyper-parameters** used in the original service, eliminating configuration drift across distributed environments.

## The Three-Step Synchronization Pipeline

The synchronization process implemented in [`phoenix/xrex/train/trainer_recsys.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_recsys.py) consists of three tightly-coupled steps that guarantee configuration consistency across distributed training jobs.

### Loading the Production Config from Checkpoint

When a training job starts or resumes from a checkpoint, the trainer locates and reads the [`config.json`](https://github.com/xai-org/x-algorithm/blob/main/config.json) file written by the production service. This occurs in the `_read_checkpoint_kafka_config` helper method inside [`phoenix/xrex/train/trainer_recsys.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_recsys.py).

```python

# phoenix/xrex/train/trainer_recsys.py

def _read_checkpoint_kafka_config(ctx) -> tuple[str | None, int | None]:
    config_path = os.path.join(ctx.checkpoint.path, _CONFIG_FILENAME)
    if not os.path.isfile(config_path):
        return None, None
    with open(config_path) as f:
        config = json.load(f)

    # The file may contain a nested "dataset" section; pull it out.

    dataset_config = config.get("dataset", config)
    topic = dataset_config.get("topic_name")
    partitions = dataset_config.get("num_kafka_partitions")
    return topic, partitions

```

The JSON file contains the full `xai_configlib` configuration hierarchy, including dataset, model, and optimizer settings.

### Deserializing with `xai_configlib`

After reading the raw JSON, the system deserializes it into a **typed configuration object** using `xai_configlib.Config.from_json`. This materializes a Python class decorated with `@configclass`, ensuring that the same structure definitions used in production are enforced during experimentation.

```python

# phoenix/xrex/train/trainer_recsys.py (excerpt)

from xai_configlib import Config
from xrex.utils.utils import broadcast_one_to_all

# After reading config.json …

raw_cfg = json.load(open(config_path))
prod_cfg = Config.from_json(raw_cfg)                # Typed config object

```

Because the codebase shares `configclass` definitions between production and experiment environments, the structures are guaranteed to match, preventing type mismatches or missing fields.

### Broadcasting to All Hosts

In multi-host training scenarios, the configuration must be identical on every process. After the root process deserializes the configuration, it propagates the object to all participants using `broadcast_one_to_all`.

```python

# Continued from previous excerpt

prod_cfg = broadcast_one_to_all(prod_cfg)          # Sync across hosts

```

Under the hood, [`phoenix/xrex/utils/utils.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/utils/utils.py) wraps `jax.experimental.multihost_utils.broadcast_one_to_all` to handle the cross-host communication. This operation completes before any training step begins, ensuring every host sees the exact same hyper-parameters.

## Persisting Configuration for Reproducibility

The metadata subsystem continuously records the active configuration throughout training, storing a JSON copy in the checkpoint directory. This creates a **single source of truth** for both production runs and experimental reruns.

In [`phoenix/xrex/utils/metadata.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/utils/metadata.py), the `record_run` method writes the configuration on every checkpoint:

```python

# phoenix/xrex/utils/metadata.py

def record_run(self, run: Run, config: Jsonable):
    """Write a copy of the config to the checkpoint directory."""
    with (path / "config.json").open("w") as f:
        f.write(self.current_config.to_json())

```

This persistence mechanism makes it trivial to resume jobs with identical parameters and audit the exact configuration used for any given checkpoint.

## Summary

- **Automatic synchronization** ensures experiment parameters match production configs by loading [`config.json`](https://github.com/xai-org/x-algorithm/blob/main/config.json) from checkpoints in [`phoenix/xrex/train/trainer_recsys.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_recsys.py).
- **Type-safe deserialization** uses `xai_configlib.Config.from_json` to materialize typed configuration objects that share schema definitions with production.
- **Distributed consistency** is achieved via `broadcast_one_to_all` in [`phoenix/xrex/utils/utils.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/utils/utils.py), which copies the config to all training hosts before execution.
- **Reproducibility** is guaranteed by the metadata system in [`phoenix/xrex/utils/metadata.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/utils/metadata.py), which persists configurations alongside checkpoints.

## Frequently Asked Questions

### How does X-Algorithm ensure all training hosts use identical parameters?

The system uses `multihost_utils.broadcast_one_to_all` (wrapped by `xrex.utils.utils.broadcast_one_to_all`) to propagate the deserialized configuration from the root process to all distributed hosts immediately after loading. This broadcast occurs before training begins, ensuring every participant starts with the exact same hyper-parameters.

### What file format is used to store production configuration?

Configurations are stored as JSON files named [`config.json`](https://github.com/xai-org/x-algorithm/blob/main/config.json) located within the checkpoint directory. This file contains the complete `xai_configlib` hierarchy including dataset, model, and optimizer settings, written by the production service when the checkpoint is created.

### Where is the configuration synchronization logic implemented?

The primary logic resides in [`phoenix/xrex/train/trainer_recsys.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_recsys.py), specifically within the `_read_checkpoint_kafka_config` method for loading and the main training setup code for deserializing and broadcasting. Supporting utilities for multi-host broadcasting live in [`phoenix/xrex/utils/utils.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/utils/utils.py).

### How does `xai_configlib` maintain type safety during config synchronization?

The library uses Python classes decorated with `@configclass` to define configuration schemas. When `Config.from_json` deserializes the checkpoint's JSON, it materializes instances of these typed classes rather than raw dictionaries. Because production and experiment code share the same `configclass` definitions, the structure and types are guaranteed to match between environments.