How to Configure Distributed Training with --num_nodes and --node_rank in TRELLIS.2
TRELLIS.2 leverages PyTorch's torch.distributed backend and exposes --num_nodes and --node_rank flags in train.py to coordinate multi-machine training jobs, automatically computing global process ranks and initializing NCCL process groups across cluster nodes.
TRELLIS.2 is Microsoft's open-source framework for 3D controllable generation that supports scaling training across multiple GPUs and physical servers. When moving beyond single-machine training, you must configure distributed training with --num_nodes and --node_rank flags to specify cluster topology and node identity. These parameters work alongside --master_addr and --master_port to establish the communication fabric required for synchronized multi-node optimization.
How Distributed Training Flags Are Parsed
In [train.py](https://github.com/microsoft/TRELLIS.2/blob/main/train.py), the command-line arguments defining cluster topology are registered using Python's argparse module:
parser.add_argument('--num_nodes', type=int, default=1,
help='Number of nodes')
parser.add_argument('--node_rank', type=int, default=0,
help='Node rank')
parser.add_argument('--master_addr', type=str, default='localhost',
help='Master address for distributed training')
parser.add_argument('--master_port', type=str, default='12345',
help='Port for distributed training')
The default values assume single-node training (--num_nodes 1, --node_rank 0), but you override these when launching multi-node jobs across physical machines.
Computing Global Rank and World Size
After parsing, the main() function calculates each process's global identity and the total process count. According to lines 60-65 in train.py:
rank = cfg.node_rank * cfg.num_gpus + local_rank
world_size = cfg.num_nodes * cfg.num_gpus
if world_size > 1:
setup_dist(rank, local_rank, world_size,
cfg.master_addr, cfg.master_port)
Here, global rank uniquely identifies each process across all nodes, calculated as node_rank * num_gpus + local_rank. The world size represents the total number of processes participating in training, computed as num_nodes * num_gpus.
Initializing the Distributed Backend
The actual environment configuration occurs in [trellis2/utils/dist_utils.py](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/utils/dist_utils.py). The setup_dist() function (lines 9-16) initializes the distributed environment:
def setup_dist(rank, local_rank, world_size, master_addr, master_port):
os.environ['MASTER_ADDR'] = master_addr
os.environ['MASTER_PORT'] = master_port
os.environ['WORLD_SIZE'] = str(world_size)
os.environ['RANK'] = str(rank)
os.environ['LOCAL_RANK'] = str(local_rank)
torch.cuda.set_device(local_rank)
dist.init_process_group('nccl', rank=rank, world_size=world_size)
This function sets the required environment variables (MASTER_ADDR, MASTER_PORT, WORLD_SIZE, RANK, LOCAL_RANK), assigns the correct CUDA device via torch.cuda.set_device(), and initializes the NCCL process group for optimal NVIDIA GPU communication.
Configuration Examples
Single-Node Multi-GPU
For training on one server with multiple GPUs, maintain the default node settings and specify only the local GPU count:
python train.py \
--config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
--output_dir results/shape_vae_8gpus \
--num_gpus 8
Here --num_nodes remains at its default value of 1 and --node_rank stays at 0.
Multi-Node Multi-GPU
For clusters with multiple physical machines, specify the total node count and individual node rank. Assuming two nodes with 4 GPUs each, using IP 10.0.0.1 as the master, launch identical commands on each node with unique rank identifiers.
On the master node (rank 0):
python train.py \
--config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
--output_dir results/shape_vae_multi_node \
--num_nodes 2 \
--node_rank 0 \
--num_gpus 4 \
--master_addr 10.0.0.1 \
--master_port 12345
On the worker node (rank 1):
python train.py \
--config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
--output_dir results/shape_vae_multi_node \
--num_nodes 2 \
--node_rank 1 \
--num_gpus 4 \
--master_addr 10.0.0.1 \
--master_port 12345
Both commands use identical --num_nodes, --master_addr, and --master_port values, differing only in --node_rank to identify their position in the cluster.
Process Spawning Mechanics
TRELLIS.2 uses torch.multiprocessing.spawn to launch one process per GPU. When you specify --num_gpus 4, the launcher spawns four processes on that node, with mp.spawn automatically supplying the local_rank argument (0 through 3) to each process's main() function. This local_rank becomes the CUDA device index for that specific process.
Important Considerations for Distributed Training
When scaling across multiple nodes, ensure the master address is reachable from all machines via the specified port. All nodes must run identical code versions and configuration files. The NCCL backend requires consistent CUDA versions across the cluster and performs best with high-bandwidth inter-node networking infrastructure.
Summary
- Argument parsing:
train.pyaccepts--num_nodes,--node_rank,--master_addr, and--master_portto define cluster topology and node identity. - Rank calculation: Global rank equals
node_rank * num_gpus + local_rank, while world size equalsnum_nodes * num_gpus. - Backend initialization:
trellis2/utils/dist_utils.pysets environment variables and initializes NCCL process groups viasetup_dist(). - Multi-node setup: Launch identical commands on each node with unique
--node_rankvalues (starting at 0) and matching--master_addrconfiguration.
Frequently Asked Questions
What is the difference between node_rank and local_rank?
--node_rank identifies which physical machine the process runs on within the cluster, starting at 0 for the master node. local_rank identifies which GPU the process uses on that specific machine, assigned automatically by torch.multiprocessing.spawn when train.py spawns nprocs=cfg.num_gpus processes. The global rank combines both values to create a unique identifier across the entire distributed job.
Do I need to change master_addr for single-node training?
No. For single-node training, the default --master_addr localhost and --master_port 12345 are sufficient because all processes communicate through the local loopback interface. You only need to specify a reachable IP address or hostname when running multi-node training across separate physical machines.
How many processes does TRELLIS.2 spawn per node?
TRELLIS.2 launches exactly one process per GPU using torch.multiprocessing.spawn with nprocs=cfg.num_gpus. If you specify --num_gpus 4, the launcher creates four processes on that node, each receiving a different local_rank (0 through 3) and corresponding CUDA device via torch.cuda.set_device(local_rank).
Can I use a different distributed backend instead of NCCL?
The setup_dist() function in trellis2/utils/dist_utils.py explicitly initializes the NCCL backend via dist.init_process_group('nccl', rank=rank, world_size=world_size). While PyTorch supports other backends like Gloo or MPI, modifying this would require changing the source code in dist_utils.py, as NCCL is hardcoded for optimal NVIDIA GPU performance in TRELLIS.2.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →