# How BiRefNet Background Removal Preprocessing Works in TRELLIS.2

> Discover how BiRefNet background removal in TRELLIS.2 automatically creates alpha channel masks for precise foreground subject isolation before 3D generation.

- Repository: [Microsoft/TRELLIS.2](https://github.com/microsoft/TRELLIS.2)
- Tags: how-to-guide
- Published: 2026-08-04

---

**BiRefNet automatically removes image backgrounds in TRELLIS.2 by running a lightweight segmentation model that produces an alpha channel mask, isolating the foreground subject before 3D generation begins.**

TRELLIS.2 uses **BiRefNet background removal preprocessing** as a mandatory first step in its image-to-3D pipeline. The system leverages the `ZhengPeng7/BiRefNet` model from Hugging Face to segment foreground objects, ensuring downstream 3D reconstruction models receive clean, distraction-free inputs. This preprocessing is handled entirely within the `Trellis2ImageTo3DPipeline` class and requires no manual intervention.

## Loading and Initializing the BiRefNet Model

The BiRefNet model is encapsulated in [`trellis2/pipelines/rembg/BiRefNet.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/rembg/BiRefNet.py). The constructor downloads the pretrained weights and configures a deterministic inference pipeline.

```python

# trellis2/pipelines/rembg/BiRefNet.py

self.model = AutoModelForImageSegmentation.from_pretrained(
    model_name, trust_remote_code=True
)
self.model.eval()

```

The model is immediately set to **evaluation mode** to disable dropout and batch normalization updates. A companion transform pipeline standardizes inputs: images are resized to **1024×1024**, converted to tensors, and normalized using ImageNet mean and standard deviation values.

## Forward Pass: Generating the Alpha Mask

When invoked, the `BiRefNet.__call__` method executes a complete segmentation workflow:

1. **Preprocessing** – The input PIL image undergoes the 1024×1024 resize and tensor conversion
2. **GPU inference** – The tensor moves to `"cuda"`, is reshaped to `(1, 3, 1024, 1024)`, and passes through the network
3. **Sigmoid activation** – Raw logits convert to per-pixel foreground probabilities
4. **Mask generation** – Probabilities become a binary alpha mask, resized back to original dimensions
5. **Alpha composition** – The mask merges with the original RGB data, producing an **RGBA PIL image**

The output preserves the original resolution while adding a fully opaque or transparent alpha channel for each pixel.

## Pipeline Integration and Memory Optimization

The `Trellis2ImageTo3DPipeline.preprocess_image` method orchestrates when and how BiRefNet runs. Located in [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py), this method implements a **low-VRAM optimization pattern** that minimizes GPU memory pressure:

```python
if self.low_vram:
    self.rembg_model.to(self.device)
output = self.rembg_model(input)
if self.low_vram:
    self.rembg_model.cpu()

```

By relocating the model to CPU immediately after inference, TRELLIS.2 frees memory for the larger 3D generation stages that follow.

## Complete Preprocessing Workflow

The full `preprocess_image` sequence includes five distinct stages:

- **Alpha detection** – RGBA inputs with existing transparency skip background removal entirely
- **Conservative resize** – Large images scale down to ≤1024 px on the longest edge, preserving aspect ratio
- **BiRefNet segmentation** – RGB images receive the alpha channel treatment described above
- **Bounding-box extraction** – The pipeline detects the tight 80% alpha threshold foreground region, centers it, and crops to a square
- **Premultiplied alpha conversion** – The final RGBA image becomes a float tensor, multiplies RGB by alpha, and returns to PIL format for conditioning models

This preprocessing guarantees that **sparse-structure**, **shape-SLat**, and **tex-SLat** models receive consistently formatted foreground subjects regardless of original background complexity.

## Using BiRefNet Directly

For debugging or custom workflows, the BiRefNet class operates independently:

```python
from trellis2.pipelines.rembg.BiRefNet import BiRefNet
from PIL import Image

rembg = BiRefNet()
rembg.to("cuda")

input_img = Image.open("my_photo.jpg")
masked_img = rembg(input_img)  # Returns RGBA with foreground isolated

masked_img.save("masked.png")

```

## Full Pipeline Example

The automatic preprocessing triggers during standard pipeline execution:

```python
from PIL import Image
from trellis2.pipelines.trellis2_image_to_3d import Trellis2ImageTo3DPipeline
import torch

pipeline = Trellis2ImageTo3DPipeline.from_pretrained(
    path="microsoft/trellis2-pretrained",
    config_file="pipeline.json"
)
pipeline.to(torch.device("cuda"))

img = Image.open("my_photo.jpg")  # No alpha channel

meshes = pipeline.run(img, num_samples=1)  # BiRefNet runs automatically

meshes[0].save("output_mesh.glb")

```

The `pipeline.run` call chains through `preprocess_image`, which internally dispatches to `self.rembg_model(input)` before any 3D operations begin.

## Key Source Files

| File | Purpose |
|------|---------|
| [`trellis2/pipelines/rembg/BiRefNet.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/rembg/BiRefNet.py) | BiRefNet model wrapper with `__call__` for alpha channel generation |
| [`trellis2/pipelines/trellis2_image_to_3d.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_image_to_3d.py) | Pipeline integration with low-VRAM memory management |
| [`trellis2/pipelines/__init__.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/__init__.py) | Submodule exposure for dynamic instantiation |
| [`trellis2/pipelines/trellis2_texturing.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/trellis2_texturing.py) | Identical preprocessing path for texture generation pipeline |

## Summary

- **BiRefNet background removal preprocessing** in TRELLIS.2 uses the Hugging Face `ZhengPeng7/BiRefNet` model to segment foreground subjects
- The `BiRefNet` class in [`trellis2/pipelines/rembg/BiRefNet.py`](https://github.com/microsoft/TRELLIS.2/blob/main/trellis2/pipelines/rembg/BiRefNet.py) handles 1024×1024 inference, sigmoid activation, and alpha channel composition
- `Trellis2ImageTo3DPipeline` automatically triggers background removal for RGB inputs and skips it for existing RGBA images
- Low-VRAM mode temporarily moves the model to GPU during inference, then returns it to CPU
- Post-segmentation crops extract the tight foreground bounding box and apply premultiplied alpha before 3D generation

## Frequently Asked Questions

### What model does TRELLIS.2 use for background removal?

TRELLIS.2 uses **BiRefNet** (`ZhengPeng7/BiRefNet` on Hugging Face), a lightweight image segmentation model specifically trained for dichotomous image segmentation. The implementation wraps `AutoModelForImageSegmentation` with a custom preprocessing pipeline that standardizes inputs to 1024×1024 resolution.

### Does TRELLIS.2 background removal run automatically?

Yes. The `Trellis2ImageTo3DPipeline.preprocess_image` method automatically detects RGB inputs and routes them through BiRefNet. Images with existing alpha channels (RGBA) bypass this step entirely. No manual configuration is required for standard usage.

### How does TRELLIS.2 manage GPU memory during background removal?

The pipeline implements a **low-VRAM mode** that moves the BiRefNet model to GPU only during the forward pass. Immediately after generating the alpha mask, the model relocates to CPU memory. This pattern prevents the ~400MB segmentation model from competing with the much larger 3D diffusion models for video memory.

### Can I use the BiRefNet background removal without the full TRELLIS.2 pipeline?

Yes. Import `BiRefNet` directly from `trellis2.pipelines.rembg.BiRefNet` and call it on any PIL image. The class returns a standard RGBA PIL image suitable for any downstream application, not just TRELLIS.2's 3D generation pipeline.