How the Image Node Loads and Prepares Images for Generation in Modly

The Modly Image node downloads or loads local images, decodes them to raw pixels, normalizes values to the [-1, 1] range, and reshapes the data into a [1, 3, H, W] tensor before transferring it via shared memory to the Python backend.

The Image node serves as the entry point for visual data in Modly's node-based workflow engine. Built for Stable Diffusion pipelines, this TypeScript/Electron component bridges the gap between user-provided image sources and the PyTorch tensors that downstream nodes require. Understanding its internal pipeline helps developers debug loading issues, optimize preprocessing, or extend support for additional image formats.

Image Node Architecture Overview

The Image node follows a three-phase execution pattern inside its process() method. Unlike asynchronous data loaders, Modly implements synchronous processing to guarantee that downstream nodes receive fully-prepared tensors immediately upon node completion.

Phase 1: Input Acquisition and Protocol Detection

The node accepts a source property that can reference either local files or remote URLs. Protocol detection determines the loading strategy:

  • file:// paths: Read directly from the filesystem using Node.js fs APIs
  • http:// or https:// URLs: Downloaded via node-fetch into a temporary memory buffer

The source handler lives in src/areas/workflows/nodes/ImageNode.tsx, where the node's main execution loop orchestrates these operations.

Phase 2: Decoding and Pixel Normalization

Once raw bytes are available, control passes to src/areas/workflows/utils/imageDecoder.ts. This utility implements format-specific decoders:

Format Library Output
JPEG jpeg-js RGBA pixel matrix
PNG pngjs RGBA pixel matrix

The decoder performs three critical transformations:

  1. Alpha channel stripping — Removes the fourth channel, leaving RGB
  2. Type casting — Converts Uint8Array values to Float32Array
  3. Range normalization — Scales [0, 255] integer values to [-1, 1] floats (the standard expected by Stable Diffusion VAEs)

Phase 3: Tensor Construction and Bridge Transfer

The normalized pixel data requires reshaping before reaching PyTorch. The Image node flattens the HWC (height × width × channels) buffer into CHW format and prepends a batch dimension:


[height, width, 3] → [1, 3, height, width]

This tensor is handed to electron/main/python-bridge.ts, which implements zero-copy shared memory transfer:

  1. Writes the Float32Array to a memory-mapped file
  2. Signals the Python process via IPC
  3. The Python side reconstructs the tensor using torch.from_numpy() in api/python/torch_utils.py

Key Implementation Files and Functions

Understanding the specific functions involved helps when tracing errors or modifying behavior.

src/areas/workflows/nodes/ImageNode.tsx

The core node implementation defines the process() method that orchestrates the entire pipeline. This is where protocol detection occurs and where the decoded tensor is handed off to the bridge.

src/areas/workflows/utils/imageDecoder.ts

Contains the decodeImage() function with signature:

function decodeImage(buffer: Buffer, mimeType: string): Float32Array

Handles JPEG/PNG parsing and returns normalized RGB data ready for tensor construction.

electron/main/python-bridge.ts

Exposes sendImageTensor(tensor: ImageTensor): Promise<LatentTensor>:

interface ImageTensor {
  data: Float32Array;
  shape: [number, number, number, number]; // [1, 3, H, W]
  dtype: 'float32';
}

Manages shared memory file creation and synchronization with the Python backend.

api/python/torch_utils.py

The Python-side entry point reconstructs PyTorch tensors:

def numpy_to_tensor(array: np.ndarray, shape: Tuple[int, ...]) -> torch.Tensor:
    """Convert shared memory numpy array to GPU tensor with correct dimensions."""
    tensor = torch.from_numpy(array).view(shape)
    return tensor.cuda() if torch.cuda.is_available() else tensor

Practical Usage Example

Defining an Image Node in Workflow JSON

{
  "nodes": [
    {
      "id": "img1",
      "type": "image",
      "config": {
        "source": "file:///Users/alice/pictures/portrait.png"
      }
    },
    {
      "id": "unet1",
      "type": "unet",
      "inputs": { "latent": "img1" }
    }
  ],
  "edges": [
    { "from": "img1", "to": "unet1", "key": "latent" }
  ]
}

Programmatic Workflow Construction

import { createWorkflow } from '@/shared/stores/workflowsStore';

const wf = createWorkflow({
  nodes: [
    {
      id: 'img',
      type: 'image',
      config: { source: 'https://example.com/cat.jpg' },
    },
    {
      id: 'vae',
      type: 'vae',
      inputs: { latent: 'img' },
    },
  ],
});

// Execution triggers download → decode → normalize → tensor transfer
await wf.run();

Performance Considerations

The Image node's synchronous design carries tradeoffs developers should understand:

  • Memory overhead: Large images are held in both JavaScript Float32Array and Python shared memory simultaneously during transfer
  • Blocking behavior: Remote URL downloads complete before process() returns, which pauses workflow execution
  • Format limitations: Only JPEG and PNG are supported in the current imageDecoder.ts implementation; WebP or AVIF would require extending the decoder utilities

For high-resolution inputs (2048×2048+), consider preprocessing to match your model's expected input resolution to reduce memory pressure during the tensor transfer phase.

Summary

  • The Image node accepts file:// paths or HTTP(S) URLs through its source configuration property
  • Decoding uses jpeg-js and pngjs libraries with alpha channel removal and [-1, 1] normalization
  • Tensor reshaping converts HWC to NCHW format [1, 3, H, W] before bridge transfer
  • Zero-copy shared memory via python-bridge.ts avoids serialization overhead when handing data to PyTorch
  • All processing occurs synchronously in process() to ensure downstream nodes receive valid tensors

Frequently Asked Questions

What image formats does the Modly Image node support?

The Image node supports JPEG and PNG through dedicated decoder libraries in src/areas/workflows/utils/imageDecoder.ts. WebP, AVIF, and other formats are not implemented and would require extending the decoder utility with additional parsing libraries.

Can the Image node load images from URLs?

Yes. The node detects http:// and https:// protocols automatically, downloads the image via node-fetch, and processes it identically to local files. The download occurs synchronously during process() execution, blocking the workflow until complete.

Why does the Image node normalize pixel values to [-1, 1] instead of [0, 1]?

Stable Diffusion VAEs are trained with [-1, 1] input normalization. The imageDecoder.ts utility scales uint8 values accordingly: pixel = (value / 127.5) - 1.0. Using [0, 1] normalization would produce incorrect latents and degraded generation quality.

How does the tensor reach the Python backend without copying data?

The electron/main/python-bridge.ts implements shared memory transfer: it writes the Float32Array to a memory-mapped file and passes the file descriptor to Python. The api/python/torch_utils.py module then uses numpy.memmap and torch.from_numpy() to create a tensor view without duplicating the underlying buffer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →