# How to Configure Distributed Training with --num_nodes and --node_rank in TRELLIS.2

> Learn to configure distributed training in TRELLIS.2 using --num_nodes and --node_rank. Master multi-machine training with PyTorch's torch.distributed backend.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**TRELLIS.2 leverages PyTorch's `torch.distributed` backend and exposes `--num_nodes` and `--node_rank` flags in [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) to coordinate multi-machine training jobs, automatically computing global process ranks and initializing NCCL process groups across cluster nodes.**

TRELLIS.2 is Microsoft's open-source framework for 3D controllable generation that supports scaling training across multiple GPUs and physical servers. When moving beyond single-machine training, you must configure distributed training with `--num_nodes` and `--node_rank` flags to specify cluster topology and node identity. These parameters work alongside `--master_addr` and `--master_port` to establish the communication fabric required for synchronized multi-node optimization.

## How Distributed Training Flags Are Parsed

In [[`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py)](https://github.com/microsoft/TRELLIS.2/blob/main/train.py), the command-line arguments defining cluster topology are registered using Python's `argparse` module:

```python
parser.add_argument('--num_nodes', type=int, default=1,
                    help='Number of nodes')
parser.add_argument('--node_rank', type=int, default=0,
                    help='Node rank')
parser.add_argument('--master_addr', type=str, default='localhost',
                    help='Master address for distributed training')
parser.add_argument('--master_port', type=str, default='12345',
                    help='Port for distributed training')

```

The default values assume single-node training (`--num_nodes 1`, `--node_rank 0`), but you override these when launching multi-node jobs across physical machines.

## Computing Global Rank and World Size

After parsing, the `main()` function calculates each process's global identity and the total process count. According to lines 60-65 in [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py):

```python
rank = cfg.node_rank * cfg.num_gpus + local_rank
world_size = cfg.num_nodes * cfg.num_gpus
if world_size > 1:
    setup_dist(rank, local_rank, world_size,
               cfg.master_addr, cfg.master_port)

```

Here, **global rank** uniquely identifies each process across all nodes, calculated as `node_rank * num_gpus + local_rank`. The **world size** represents the total number of processes participating in training, computed as `num_nodes * num_gpus`.

## Initializing the Distributed Backend

The actual environment configuration occurs in [[`trellis2/utils/dist_utils.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/utils/dist_utils.py)](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/utils/dist_utils.py). The `setup_dist()` function (lines 9-16) initializes the distributed environment:

```python
def setup_dist(rank, local_rank, world_size, master_addr, master_port):
    os.environ['MASTER_ADDR'] = master_addr
    os.environ['MASTER_PORT'] = master_port
    os.environ['WORLD_SIZE'] = str(world_size)
    os.environ['RANK'] = str(rank)
    os.environ['LOCAL_RANK'] = str(local_rank)
    torch.cuda.set_device(local_rank)
    dist.init_process_group('nccl', rank=rank, world_size=world_size)

```

This function sets the required environment variables (`MASTER_ADDR`, `MASTER_PORT`, `WORLD_SIZE`, `RANK`, `LOCAL_RANK`), assigns the correct CUDA device via `torch.cuda.set_device()`, and initializes the **NCCL process group** for optimal NVIDIA GPU communication.

## Configuration Examples

### Single-Node Multi-GPU

For training on one server with multiple GPUs, maintain the default node settings and specify only the local GPU count:

```bash
python train.py \
  --config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
  --output_dir results/shape_vae_8gpus \
  --num_gpus 8

```

Here `--num_nodes` remains at its default value of `1` and `--node_rank` stays at `0`.

### Multi-Node Multi-GPU

For clusters with multiple physical machines, specify the total node count and individual node rank. Assuming two nodes with 4 GPUs each, using IP `10.0.0.1` as the master, launch identical commands on each node with unique rank identifiers.

On the **master node** (rank 0):

```bash
python train.py \
  --config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
  --output_dir results/shape_vae_multi_node \
  --num_nodes 2 \
  --node_rank 0 \
  --num_gpus 4 \
  --master_addr 10.0.0.1 \
  --master_port 12345

```

On the **worker node** (rank 1):

```bash
python train.py \
  --config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
  --output_dir results/shape_vae_multi_node \
  --num_nodes 2 \
  --node_rank 1 \
  --num_gpus 4 \
  --master_addr 10.0.0.1 \
  --master_port 12345

```

Both commands use identical `--num_nodes`, `--master_addr`, and `--master_port` values, differing only in `--node_rank` to identify their position in the cluster.

### Process Spawning Mechanics

TRELLIS.2 uses `torch.multiprocessing.spawn` to launch one process per GPU. When you specify `--num_gpus 4`, the launcher spawns four processes on that node, with `mp.spawn` automatically supplying the `local_rank` argument (0 through 3) to each process's `main()` function. This `local_rank` becomes the CUDA device index for that specific process.

## Important Considerations for Distributed Training

When scaling across multiple nodes, ensure the **master address** is reachable from all machines via the specified port. All nodes must run identical code versions and configuration files. The NCCL backend requires consistent CUDA versions across the cluster and performs best with high-bandwidth inter-node networking infrastructure.

## Summary

- **Argument parsing**: [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) accepts `--num_nodes`, `--node_rank`, `--master_addr`, and `--master_port` to define cluster topology and node identity.
- **Rank calculation**: Global rank equals `node_rank * num_gpus + local_rank`, while world size equals `num_nodes * num_gpus`.
- **Backend initialization**: [`trellis2/utils/dist_utils.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/utils/dist_utils.py) sets environment variables and initializes NCCL process groups via `setup_dist()`.
- **Multi-node setup**: Launch identical commands on each node with unique `--node_rank` values (starting at 0) and matching `--master_addr` configuration.

## Frequently Asked Questions

### What is the difference between node_rank and local_rank?

`--node_rank` identifies which physical machine the process runs on within the cluster, starting at 0 for the master node. `local_rank` identifies which GPU the process uses on that specific machine, assigned automatically by `torch.multiprocessing.spawn` when [`train.py`](https://github.com/microsoft/TRELLIS.2/blob/main/train.py) spawns `nprocs=cfg.num_gpus` processes. The global rank combines both values to create a unique identifier across the entire distributed job.

### Do I need to change master_addr for single-node training?

No. For single-node training, the default `--master_addr localhost` and `--master_port 12345` are sufficient because all processes communicate through the local loopback interface. You only need to specify a reachable IP address or hostname when running multi-node training across separate physical machines.

### How many processes does TRELLIS.2 spawn per node?

TRELLIS.2 launches exactly one process per GPU using `torch.multiprocessing.spawn` with `nprocs=cfg.num_gpus`. If you specify `--num_gpus 4`, the launcher creates four processes on that node, each receiving a different `local_rank` (0 through 3) and corresponding CUDA device via `torch.cuda.set_device(local_rank)`.

### Can I use a different distributed backend instead of NCCL?

The `setup_dist()` function in [`trellis2/utils/dist_utils.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/utils/dist_utils.py) explicitly initializes the NCCL backend via `dist.init_process_group('nccl', rank=rank, world_size=world_size)`. While PyTorch supports other backends like Gloo or MPI, modifying this would require changing the source code in [`dist_utils.py`](https://github.com/microsoft/TRELLIS.2/blob/main/dist_utils.py), as NCCL is hardcoded for optimal NVIDIA GPU performance in TRELLIS.2.