How to Use Diffusion Models for Image Generation with AILIA Models
You can generate high-quality images from text prompts using the ready-to-run Stable Diffusion and SDXL-Turbo pipelines in the axinc-ai/ailia-models repository, which handle text encoding, iterative denoising, and VAE decoding automatically.
The axinc-ai/ailia-models repository provides complete, production-ready implementations of state-of-the-art diffusion models for image generation. These pipelines encapsulate the full generation workflow—from CLIP text encoding through UNet denoising to VAE decoding—allowing you to create images via simple command-line interfaces or direct Python integration.
Understanding the Diffusion Pipeline Architecture
The diffusion pipelines in this repository follow the standard latent diffusion architecture established by Stable Diffusion and SDXL-Turbo. Each component is implemented as a distinct ONNX model loaded through the AILIA SDK or ONNX Runtime.
Core Components
The pipelines consist of five interconnected components:
-
Text Encoder (CLIP): Converts natural language prompts into latent text embeddings. Implemented in
stable-diffusion-txt2img.pyviaFrozenCLIPEmbedderand insdxl-turbo.pythrough dual text encoders (lines 13-31). -
UNet (Denoising Network): Predicts noise residuals for latent images at specific timesteps, conditioned on text embeddings. Loaded via
ailia.Netoronnxruntime.InferenceSessionusing paths likeWEIGHT_UNET_PATHinsdxl-turbo.py. -
VAE Encoder/Decoder: Transforms between pixel space and the lower-dimensional latent space where diffusion occurs. The decoder (
vae_decoder) converts final latents to RGB images in both pipelines. -
Scheduler: Manages timestep scheduling and reverse diffusion steps. Uses
EulerAncestralDiscreteSchedulerin SDXL-Turbo and custom DDIM helpers (make_ddim_timesteps,make_ddim_sampling_parameters) instable-diffusion-txt2img.py. -
Model Wrappers: High-level classes (
StableDiffusionXL,StableDiffusionXLImg2Imginsdxl-turbo/df/pipelines/) that orchestrate component interactions.
The Generation Process
The diffusion process follows these discrete steps:
- Encode the prompt using CLIP to produce text embedding
c. - Sample random latent
zfrom Gaussian noise with shape[C, H/f, W/f]. - Iteratively denoise
zfor N timesteps using the UNet, scheduler, and classifier-free guidance scale. - Decode the final latent through the VAE decoder to produce the RGB output image.
Running Stable Diffusion for Text-to-Image Generation
The stable-diffusion-txt2img.py script provides a complete CLI for generating images from text prompts using Stable Diffusion v1.4 or custom checkpoints.
Installation Requirements
Install the required dependency for text tokenization:
pip3 install transformers
Basic CLI Usage
Generate an image with a single command:
python3 diffusion/stable-diffusion-txt2img/stable-diffusion-txt2img.py \
--input "a photograph of an astronaut riding a horse" \
--savepath results/astro.png
The script automatically downloads required ONNX models from the remote bucket on first run.
Common Configuration Options
| Flag | Description |
|---|---|
--n_iter N |
Repeat generation N times (creates result grid) |
--n_samples S |
Generate S samples per iteration |
--steps K |
Number of DDIM sampling steps (default: 50) |
--scale V |
Classifier-free guidance scale (default: 7.5) |
--sampler {PLMS,DDIM,"DPM++ 2M Kerras"} |
Sampling algorithm selection |
--onnx |
Execute via ONNX Runtime instead of AILIA SDK |
--legacy |
Use three-file legacy model layout (diffusion_emb, diffusion_mid, diffusion_out) |
--sd {default,basil_mix} |
Switch to Basil-Mix checkpoint |
--vae {default,vae-ft-mse} |
Use fine-tuned VAE for higher quality |
Using SDXL-Turbo for Faster Inference
The sdxl-turbo.py script implements the SDXL-Turbo architecture, which generates high-quality images in as few as 1-4 steps compared to the 50 steps required by standard Stable Diffusion.
Text-to-Image Generation
Run SDXL-Turbo with minimal configuration:
python3 diffusion/sdxl-turbo/sdxl-turbo.py \
--input "little cute gremlin wearing a jacket, cinematic, vivid colors" \
--savepath results/gremlin.png
Image-to-Image Generation
Supply an initial image to guide the generation process:
python3 diffusion/sdxl-turbo/sdxl-turbo.py \
--input "cat wizard, Gandalf, fantasy, Pixar style" \
--init_image samples/cat.png \
--savepath results/cat_wizard.png
SDXL-Turbo Specific Options
| Flag | Description |
|---|---|
--init_image PATH |
Source image for img2img pipeline |
--seed N |
Fixed random seed for reproducible outputs |
--disable_ailia_tokenizer |
Use Hugging Face tokenizer instead of AILIA's bundled version |
The SDXL-Turbo implementation uses dual text encoders and the EulerAncestralDiscreteScheduler defined in diffusion/sdxl-turbo/df/schedulers/scheduling_euler_ancestral_discrete.py.
Integrating Diffusion Models Programmatically
Beyond CLI usage, you can import the pipeline components directly into your Python applications.
Importing from Stable Diffusion
Access helper functions and the main entry point from stable-diffusion-txt2img.py:
from diffusion.stable_diffusion_txt2img.stable_diffusion_txt2img import (
make_ddim_timesteps,
make_ddim_sampling_parameters,
recognize_from_text,
main as sd_main,
)
# Run generation without CLI parsing
sd_main()
Using the StableDiffusionXL Class
For SDXL-Turbo, instantiate the StableDiffusionXL class directly from the df sub-package:
from diffusion.sdxl_turbo.df import StableDiffusionXL
pipe = StableDiffusionXL(
vae_decoder=vae_decoder,
text_encoder=text_encoder,
text_encoder_2=text_encoder_2,
tokenizer=tokenizer,
tokenizer_2=tokenizer_2,
unet=unet,
scheduler=scheduler,
use_onnx=False,
)
image = pipe.forward(
prompt="a futuristic city at sunset",
num_inference_steps=2
)
The class definition resides in diffusion/sdxl-turbo/df/pipelines/stable_diffusion_xl.py, with the img2img variant available in stable_diffusion_xl_img2img.py.
Summary
- The axinc-ai/ailia-models repository provides complete diffusion pipelines for both standard Stable Diffusion and high-speed SDXL-Turbo generation.
- Each pipeline consists of five core components: CLIP text encoder, UNet denoiser, VAE encoder/decoder, scheduler, and model wrappers.
- Run text-to-image generation via
stable-diffusion-txt2img.pyorsdxl-turbo.pywith automatic model downloading and support for both AILIA SDK and ONNX Runtime backends. - Implement image-to-image generation using the
--init_imageflag in SDXL-Turbo or programmatically via theStableDiffusionXLImg2Imgclass. - Integrate diffusion capabilities into custom applications by importing helper functions from the scripts or instantiating the
StableDiffusionXLclass directly.
Frequently Asked Questions
What is the difference between Stable Diffusion and SDXL-Turbo in the ailia-models repository?
Stable Diffusion (v1.4/Basil-Mix) requires approximately 50 inference steps to generate high-quality images and uses a single text encoder. SDXL-Turbo generates comparable quality in 1-4 steps using a distilled UNet and dual text encoders (CLIP-L and OpenCLIP-G), making it significantly faster for real-time applications. Both are implemented in the repository with identical CLI interfaces but reside in different directories (stable-diffusion-txt2img vs sdxl-turbo).
Do I need the AILIA SDK to run these diffusion models?
No. While the scripts default to the AILIA SDK for optimized inference, you can run both pipelines entirely with ONNX Runtime by adding the --onnx flag to any command. This allows execution on CPU or CUDA without installing the proprietary AILIA SDK, though performance characteristics may differ depending on your hardware and ONNX Runtime version.
How do I perform image-to-image generation with these models?
For SDXL-Turbo, use the --init_image flag followed by the path to your source image: python3 diffusion/sdxl-turbo/sdxl-turbo.py --input "prompt" --init_image source.jpg. Programmatically, import StableDiffusionXLImg2Img from diffusion.sdxl_turbo.df and pass the image parameter to the forward() method. The standard Stable Diffusion pipeline in the repository focuses on text-to-image, though the underlying architecture supports img2img via the VAE encoder component.
What hardware requirements are needed for reasonable inference speeds?
Both pipelines run on CPU via ONNX Runtime, but a CUDA-capable GPU with at least 8GB VRAM is recommended for production use. SDXL-Turbo specifically benefits from GPU acceleration due to its larger UNet and dual text encoders. The AILIA SDK provides optimized backends for various edge devices including ARM processors and NVIDIA Jetson platforms, making these models deployable on embedded systems with reduced precision (FP16/INT8) where supported.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →