# How the Image Node Loads and Prepares Images for Generation in Modly

> Learn how the Modly Image node loads, decodes, normalizes, and reshapes images into tensors for AI generation, preparing them efficiently for the Python backend.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: internals
- Published: 2026-08-20

---

**The Modly Image node downloads or loads local images, decodes them to raw pixels, normalizes values to the [-1, 1] range, and reshapes the data into a [1, 3, H, W] tensor before transferring it via shared memory to the Python backend.**

The **Image node** serves as the entry point for visual data in Modly's node-based workflow engine. Built for Stable Diffusion pipelines, this TypeScript/Electron component bridges the gap between user-provided image sources and the PyTorch tensors that downstream nodes require. Understanding its internal pipeline helps developers debug loading issues, optimize preprocessing, or extend support for additional image formats.

## Image Node Architecture Overview

The Image node follows a three-phase execution pattern inside its `process()` method. Unlike asynchronous data loaders, Modly implements **synchronous processing** to guarantee that downstream nodes receive fully-prepared tensors immediately upon node completion.

### Phase 1: Input Acquisition and Protocol Detection

The node accepts a `source` property that can reference either local files or remote URLs. Protocol detection determines the loading strategy:

- **`file://` paths**: Read directly from the filesystem using Node.js `fs` APIs
- **`http://` or `https://` URLs**: Downloaded via `node-fetch` into a temporary memory buffer

The source handler lives in [`src/areas/workflows/nodes/ImageNode.tsx`](https://github.com/lightningpixel/modly/blob/main/src/areas/workflows/nodes/ImageNode.tsx), where the node's main execution loop orchestrates these operations.

### Phase 2: Decoding and Pixel Normalization

Once raw bytes are available, control passes to [`src/areas/workflows/utils/imageDecoder.ts`](https://github.com/lightningpixel/modly/blob/main/src/areas/workflows/utils/imageDecoder.ts). This utility implements format-specific decoders:

| Format | Library | Output |
|--------|---------|--------|
| JPEG | `jpeg-js` | RGBA pixel matrix |
| PNG | `pngjs` | RGBA pixel matrix |

The decoder performs three critical transformations:

1. **Alpha channel stripping** — Removes the fourth channel, leaving RGB
2. **Type casting** — Converts `Uint8Array` values to `Float32Array`
3. **Range normalization** — Scales [0, 255] integer values to [-1, 1] floats (the standard expected by Stable Diffusion VAEs)

### Phase 3: Tensor Construction and Bridge Transfer

The normalized pixel data requires reshaping before reaching PyTorch. The Image node flattens the HWC (height × width × channels) buffer into CHW format and prepends a batch dimension:

```

[height, width, 3] → [1, 3, height, width]

```

This tensor is handed to [`electron/main/python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/python-bridge.ts), which implements **zero-copy shared memory transfer**:

1. Writes the `Float32Array` to a memory-mapped file
2. Signals the Python process via IPC
3. The Python side reconstructs the tensor using `torch.from_numpy()` in [`api/python/torch_utils.py`](https://github.com/lightningpixel/modly/blob/main/api/python/torch_utils.py)

## Key Implementation Files and Functions

Understanding the specific functions involved helps when tracing errors or modifying behavior.

### src/areas/workflows/nodes/ImageNode.tsx

The core node implementation defines the `process()` method that orchestrates the entire pipeline. This is where protocol detection occurs and where the decoded tensor is handed off to the bridge.

### src/areas/workflows/utils/imageDecoder.ts

Contains the `decodeImage()` function with signature:

```typescript
function decodeImage(buffer: Buffer, mimeType: string): Float32Array

```

Handles JPEG/PNG parsing and returns normalized RGB data ready for tensor construction.

### electron/main/python-bridge.ts

Exposes `sendImageTensor(tensor: ImageTensor): Promise<LatentTensor>`:

```typescript
interface ImageTensor {
  data: Float32Array;
  shape: [number, number, number, number]; // [1, 3, H, W]
  dtype: 'float32';
}

```

Manages shared memory file creation and synchronization with the Python backend.

### api/python/torch_utils.py

The Python-side entry point reconstructs PyTorch tensors:

```python
def numpy_to_tensor(array: np.ndarray, shape: Tuple[int, ...]) -> torch.Tensor:
    """Convert shared memory numpy array to GPU tensor with correct dimensions."""
    tensor = torch.from_numpy(array).view(shape)
    return tensor.cuda() if torch.cuda.is_available() else tensor

```

## Practical Usage Example

### Defining an Image Node in Workflow JSON

```json
{
  "nodes": [
    {
      "id": "img1",
      "type": "image",
      "config": {
        "source": "file:///Users/alice/pictures/portrait.png"
      }
    },
    {
      "id": "unet1",
      "type": "unet",
      "inputs": { "latent": "img1" }
    }
  ],
  "edges": [
    { "from": "img1", "to": "unet1", "key": "latent" }
  ]
}

```

### Programmatic Workflow Construction

```typescript
import { createWorkflow } from '@/shared/stores/workflowsStore';

const wf = createWorkflow({
  nodes: [
    {
      id: 'img',
      type: 'image',
      config: { source: 'https://example.com/cat.jpg' },
    },
    {
      id: 'vae',
      type: 'vae',
      inputs: { latent: 'img' },
    },
  ],
});

// Execution triggers download → decode → normalize → tensor transfer
await wf.run();

```

## Performance Considerations

The Image node's synchronous design carries tradeoffs developers should understand:

- **Memory overhead**: Large images are held in both JavaScript `Float32Array` and Python shared memory simultaneously during transfer
- **Blocking behavior**: Remote URL downloads complete before `process()` returns, which pauses workflow execution
- **Format limitations**: Only JPEG and PNG are supported in the current [`imageDecoder.ts`](https://github.com/lightningpixel/modly/blob/main/imageDecoder.ts) implementation; WebP or AVIF would require extending the decoder utilities

For high-resolution inputs (2048×2048+), consider preprocessing to match your model's expected input resolution to reduce memory pressure during the tensor transfer phase.

## Summary

- The **Image node** accepts `file://` paths or HTTP(S) URLs through its `source` configuration property
- **Decoding** uses `jpeg-js` and `pngjs` libraries with alpha channel removal and [-1, 1] normalization
- **Tensor reshaping** converts HWC to NCHW format [1, 3, H, W] before bridge transfer
- **Zero-copy shared memory** via [`python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/python-bridge.ts) avoids serialization overhead when handing data to PyTorch
- All processing occurs **synchronously** in `process()` to ensure downstream nodes receive valid tensors

## Frequently Asked Questions

### What image formats does the Modly Image node support?

The Image node supports **JPEG and PNG** through dedicated decoder libraries in [`src/areas/workflows/utils/imageDecoder.ts`](https://github.com/lightningpixel/modly/blob/main/src/areas/workflows/utils/imageDecoder.ts). WebP, AVIF, and other formats are not implemented and would require extending the decoder utility with additional parsing libraries.

### Can the Image node load images from URLs?

Yes. The node detects `http://` and `https://` protocols automatically, downloads the image via `node-fetch`, and processes it identically to local files. The download occurs synchronously during `process()` execution, blocking the workflow until complete.

### Why does the Image node normalize pixel values to [-1, 1] instead of [0, 1]?

Stable Diffusion VAEs are trained with **[-1, 1] input normalization**. The [`imageDecoder.ts`](https://github.com/lightningpixel/modly/blob/main/imageDecoder.ts) utility scales uint8 values accordingly: `pixel = (value / 127.5) - 1.0`. Using [0, 1] normalization would produce incorrect latents and degraded generation quality.

### How does the tensor reach the Python backend without copying data?

The [`electron/main/python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/python-bridge.ts) implements **shared memory transfer**: it writes the `Float32Array` to a memory-mapped file and passes the file descriptor to Python. The [`api/python/torch_utils.py`](https://github.com/lightningpixel/modly/blob/main/api/python/torch_utils.py) module then uses `numpy.memmap` and `torch.from_numpy()` to create a tensor view without duplicating the underlying buffer.