How the LuisaCompute `Tensor` Class Implements 1D, 2D, and 3D Convolutions with Padding

The Tensor class in luisagroup/luisacompute handles convolution padding explicitly through a private pad helper before executing dimension-specific loops in src/tensor/expression.cpp to compute conv_1d, conv_2d, and conv_3d operations.

The luisacompute repository provides a high-performance tensor computation backend that exposes unified convolution primitives across multiple dimensions. These operations treat padding as a discrete preprocessing step, ensuring boundary conditions are resolved before the kernel slides over spatial dimensions.

Architecture of the Convolution Pipeline

All three convolution variants—conv_1d, conv_2d, and conv_3d—share a common execution pipeline defined in src/tensor/expression.cpp. The implementation separates spatial boundary handling from the core dot-product computation, allowing the same logical flow to adapt to arbitrary dimensionality.

The Padding Preprocessing Step

Before any convolution arithmetic occurs, each public method invokes the private helper Tensor::pad. This helper accepts a Padding struct that specifies how many elements to prepend and append on every spatial dimension. The routine allocates a new Tensor with extents expanded by the padding offsets, then copies the original data into the interior region. This creates a zero-padded (or custom-padded) buffer that eliminates boundary checks during the subsequent convolution loop.

Kernel Preparation and Dilation

The implementation expects the kernel tensor to arrive pre-shaped as [out_channels, in_channels, *kernel_dims]. No additional padding is applied to the kernel itself, but the routines accept a dilation parameter that expands the effective receptive field without increasing the physical kernel size. The dilation factor is folded into the stride calculations during loop indexing.

Dimension-Specific Execution Loops

After padding, the three methods diverge into dimension-specific nested loops that iterate over the output spatial extents.

1D Convolution Mechanics

In conv_1d, the implementation allocates an output tensor sized according to the padded input length, kernel size, and stride parameters. For each output position, it computes a dot-product between the 1D kernel slice and the corresponding window of the padded input tensor.

2D and 3D Spatial Sliding

conv_2d and conv_3d extend this pattern with two and three nested spatial loops respectively. Each iteration extracts a multi-dimensional slice from the padded buffer and performs a dot-product against the kernel. The stride arguments—provided as std::array<int, 2> for 2D and std::array<int, 3> for 3D—control the step size across height, width, and depth dimensions.

API Signatures and Source Locations

The public interface declared in include/luisa/tensor.h and implemented in src/tensor/expression.cpp (approximately lines 140–190) exposes the following methods:

// 1D convolution with explicit padding
Tensor Tensor::conv_1d(const Tensor &kernel,
                       const std::array<int, 2> &stride,
                       const Padding &padding);

// 2D convolution with explicit padding
Tensor Tensor::conv_2d(const Tensor &kernel,
                       const std::array<int, 2> &stride,
                       const Padding &padding);

// 3D convolution with explicit padding
Tensor Tensor::conv_3d(const Tensor &kernel,
                       const std::array<int, 3> &stride,
                       const Padding &padding);

All three signatures forward to shared internal logic that manages memory layout and dispatch, differing only in the dimensionality of the traversal loops.

Practical Code Examples

C++ Implementation with Zero Padding

The following example demonstrates a 2D convolution with a 1-pixel border padding applied to height and width dimensions:

#include <luisa/compute.h>
using namespace luisa::compute;

int main() {
    // Input: 1 batch, 3 channels, 32×32 spatial
    Tensor input = Tensor::zeros({1, 3, 32, 32});

    // Kernel: 16 output channels, 3 input channels, 3×3 spatial
    Tensor kernel = Tensor::randn({16, 3, 3, 3});

    // Define padding: no padding on batch/channel, 1 pixel on H and W
    Tensor::Padding pad = {{0, 0},   // batch dimension
                           {0, 0},   // channel dimension
                           {1, 1},   // height (top, bottom)
                           {1, 1}};  // width (left, right)

    std::array<int, 2> stride = {1, 1};

    // Execute convolution; output shape is {1, 16, 32, 32}
    Tensor output = input.conv_2d(kernel, stride, pad);
}

Python Interoperability

Through the PyTorch interop layer tested in src/tests/python/test-pytorch-interop.py, the same operations are accessible from Python:

import torch
import luisa

# Create torch tensors and wrap as Luisa Tensors

x = torch.randn(1, 3, 32, 32).cuda()
lx = luisa.from_torch(x)

k = torch.randn(16, 3, 3, 3).cuda()
lk = luisa.from_torch(k)

# Padding and stride configuration

padding = luisa.Padding(((0,0), (0,0), (1,1), (1,1)))
stride = (1, 1)

# Perform convolution via the Tensor class

y = lx.conv_2d(lk, stride, padding)
torch_y = y.to_torch()

Summary

  • The Tensor::pad helper in src/tensor/expression.cpp preprocesses spatial boundaries before convolution execution.
  • conv_1d, conv_2d, and conv_3d share core logic but implement dimension-specific nested loops for sliding window operations.
  • The kernel must be shaped as [out_channels, in_channels, *kernel_dims] with no internal padding.
  • Padding is explicitly defined per dimension through a structured configuration object, enabling asymmetric or zero-padding scenarios.
  • Both C++ native code and Python PyTorch interoperability layers consume the same underlying C++ implementation.

Frequently Asked Questions

How does the Padding struct define spatial padding in the Tensor class?

The Padding struct accepts an array of integer pairs representing the number of elements to prepend and append for each tensor dimension. For a 2D convolution, you typically provide four pairs: batch (usually [0, 0]), channels ([0, 0]), height ([top, bottom]), and width ([left, right]). This explicit specification allows asymmetric padding configurations that are applied by the private Tensor::pad method before the convolution kernel begins execution.

What tensor format does the kernel parameter expect?

The kernel tensor must follow the shape convention [out_channels, in_channels, *kernel_dims] where kernel_dims matches the spatial dimensionality of the operation (length for 1D, height×width for 2D, depth×height×width for 3D). The implementation in src/tensor/expression.cpp assumes this layout to correctly index the filter weights during the dot-product calculation against the padded input slices.

Does the Tensor class support dilated convolutions?

Yes. While the primary parameters specify stride and padding, the convolution routines accept a dilation parameter that modifies the effective spacing between kernel elements. The dilation factor is incorporated into the indexing arithmetic within the convolution loops, allowing expanded receptive fields without increasing the physical kernel buffer size or reallocating the filter tensor.

Where is the core convolution logic implemented?

The core logic resides in src/tensor/expression.cpp (specifically around lines 140–190 according to repository analysis), where conv_1d, conv_2d, and conv_3d are defined alongside the private pad helper. The public API is declared in include/luisa/tensor.h, and practical usage is demonstrated in src/tests/python/test-pytorch-interop.py and src/tests/test_mnist.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →