# How the LuisaCompute `Tensor` Class Implements 1D, 2D, and 3D Convolutions with Padding

> Discover how the LuisaCompute Tensor class expertly handles conv_1d, conv_2d, and conv_3d operations with padding by examining its private pad helper and dimension-specific loops in expression.cpp.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

**The `Tensor` class in luisagroup/luisacompute handles convolution padding explicitly through a private `pad` helper before executing dimension-specific loops in [`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp) to compute `conv_1d`, `conv_2d`, and `conv_3d` operations.**

The **luisacompute** repository provides a high-performance tensor computation backend that exposes unified convolution primitives across multiple dimensions. These operations treat padding as a discrete preprocessing step, ensuring boundary conditions are resolved before the kernel slides over spatial dimensions.

## Architecture of the Convolution Pipeline

All three convolution variants—`conv_1d`, `conv_2d`, and `conv_3d`—share a common execution pipeline defined in **[`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp)**. The implementation separates spatial boundary handling from the core dot-product computation, allowing the same logical flow to adapt to arbitrary dimensionality.

### The Padding Preprocessing Step

Before any convolution arithmetic occurs, each public method invokes the private helper **`Tensor::pad`**. This helper accepts a **`Padding`** struct that specifies how many elements to prepend and append on every spatial dimension. The routine allocates a new `Tensor` with extents expanded by the padding offsets, then copies the original data into the interior region. This creates a zero-padded (or custom-padded) buffer that eliminates boundary checks during the subsequent convolution loop.

### Kernel Preparation and Dilation

The implementation expects the kernel tensor to arrive pre-shaped as **`[out_channels, in_channels, *kernel_dims]`**. No additional padding is applied to the kernel itself, but the routines accept a **dilation** parameter that expands the effective receptive field without increasing the physical kernel size. The dilation factor is folded into the stride calculations during loop indexing.

## Dimension-Specific Execution Loops

After padding, the three methods diverge into dimension-specific nested loops that iterate over the output spatial extents.

### 1D Convolution Mechanics

In **`conv_1d`**, the implementation allocates an output tensor sized according to the padded input length, kernel size, and stride parameters. For each output position, it computes a dot-product between the 1D kernel slice and the corresponding window of the padded input tensor.

### 2D and 3D Spatial Sliding

**`conv_2d`** and **`conv_3d`** extend this pattern with two and three nested spatial loops respectively. Each iteration extracts a multi-dimensional slice from the padded buffer and performs a dot-product against the kernel. The stride arguments—provided as `std::array<int, 2>` for 2D and `std::array<int, 3>` for 3D—control the step size across height, width, and depth dimensions.

## API Signatures and Source Locations

The public interface declared in **[`include/luisa/tensor.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/tensor.h)** and implemented in **[`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp)** (approximately lines 140–190) exposes the following methods:

```cpp
// 1D convolution with explicit padding
Tensor Tensor::conv_1d(const Tensor &kernel,
                       const std::array<int, 2> &stride,
                       const Padding &padding);

// 2D convolution with explicit padding
Tensor Tensor::conv_2d(const Tensor &kernel,
                       const std::array<int, 2> &stride,
                       const Padding &padding);

// 3D convolution with explicit padding
Tensor Tensor::conv_3d(const Tensor &kernel,
                       const std::array<int, 3> &stride,
                       const Padding &padding);

```

All three signatures forward to shared internal logic that manages memory layout and dispatch, differing only in the dimensionality of the traversal loops.

## Practical Code Examples

### C++ Implementation with Zero Padding

The following example demonstrates a 2D convolution with a 1-pixel border padding applied to height and width dimensions:

```cpp
#include <luisa/compute.h>
using namespace luisa::compute;

int main() {
    // Input: 1 batch, 3 channels, 32×32 spatial
    Tensor input = Tensor::zeros({1, 3, 32, 32});

    // Kernel: 16 output channels, 3 input channels, 3×3 spatial
    Tensor kernel = Tensor::randn({16, 3, 3, 3});

    // Define padding: no padding on batch/channel, 1 pixel on H and W
    Tensor::Padding pad = {{0, 0},   // batch dimension
                           {0, 0},   // channel dimension
                           {1, 1},   // height (top, bottom)
                           {1, 1}};  // width (left, right)

    std::array<int, 2> stride = {1, 1};

    // Execute convolution; output shape is {1, 16, 32, 32}
    Tensor output = input.conv_2d(kernel, stride, pad);
}

```

### Python Interoperability

Through the PyTorch interop layer tested in **[`src/tests/python/test-pytorch-interop.py`](https://github.com/luisagroup/luisacompute/blob/main/src/tests/python/test-pytorch-interop.py)**, the same operations are accessible from Python:

```python
import torch
import luisa

# Create torch tensors and wrap as Luisa Tensors

x = torch.randn(1, 3, 32, 32).cuda()
lx = luisa.from_torch(x)

k = torch.randn(16, 3, 3, 3).cuda()
lk = luisa.from_torch(k)

# Padding and stride configuration

padding = luisa.Padding(((0,0), (0,0), (1,1), (1,1)))
stride = (1, 1)

# Perform convolution via the Tensor class

y = lx.conv_2d(lk, stride, padding)
torch_y = y.to_torch()

```

## Summary

- The **`Tensor::pad`** helper in [`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp) preprocesses spatial boundaries before convolution execution.
- **`conv_1d`**, **`conv_2d`**, and **`conv_3d`** share core logic but implement dimension-specific nested loops for sliding window operations.
- The **kernel** must be shaped as `[out_channels, in_channels, *kernel_dims]` with no internal padding.
- **Padding** is explicitly defined per dimension through a structured configuration object, enabling asymmetric or zero-padding scenarios.
- Both C++ native code and Python PyTorch interoperability layers consume the same underlying C++ implementation.

## Frequently Asked Questions

### How does the Padding struct define spatial padding in the Tensor class?

The **Padding** struct accepts an array of integer pairs representing the number of elements to prepend and append for each tensor dimension. For a 2D convolution, you typically provide four pairs: batch (usually `[0, 0]`), channels (`[0, 0]`), height (`[top, bottom]`), and width (`[left, right]`). This explicit specification allows asymmetric padding configurations that are applied by the private `Tensor::pad` method before the convolution kernel begins execution.

### What tensor format does the kernel parameter expect?

The **kernel** tensor must follow the shape convention `[out_channels, in_channels, *kernel_dims]` where `kernel_dims` matches the spatial dimensionality of the operation (length for 1D, height×width for 2D, depth×height×width for 3D). The implementation in [`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp) assumes this layout to correctly index the filter weights during the dot-product calculation against the padded input slices.

### Does the Tensor class support dilated convolutions?

Yes. While the primary parameters specify stride and padding, the convolution routines accept a **dilation** parameter that modifies the effective spacing between kernel elements. The dilation factor is incorporated into the indexing arithmetic within the convolution loops, allowing expanded receptive fields without increasing the physical kernel buffer size or reallocating the filter tensor.

### Where is the core convolution logic implemented?

The core logic resides in **[`src/tensor/expression.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tensor/expression.cpp)** (specifically around lines 140–190 according to repository analysis), where `conv_1d`, `conv_2d`, and `conv_3d` are defined alongside the private `pad` helper. The public API is declared in **[`include/luisa/tensor.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/tensor.h)**, and practical usage is demonstrated in **[`src/tests/python/test-pytorch-interop.py`](https://github.com/luisagroup/luisacompute/blob/main/src/tests/python/test-pytorch-interop.py)** and **[`src/tests/test_mnist.py`](https://github.com/luisagroup/luisacompute/blob/main/src/tests/test_mnist.py)**.