# How do Multi-Modal LLMs Process and Generate Text and Images: A Deep Dive into NExT-GPT

> Discover how multi-modal LLMs process and generate text and images. Learn about modality-specific encoders, joint reasoning, and specialized decoders like Stable Diffusion.

- Repository: [Tongxin Yuan/dive-into-llms](https://github.com/Lordog/dive-into-llms)
- Tags: deep-dive
- Published: 2026-04-16

---

**Multi-modal LLMs process text and images by using modality-specific encoders to convert inputs into embeddings, projecting these into the LLM's hidden space for joint reasoning, and routing outputs through specialized decoders—such as Stable Diffusion for images or the language head for text—to generate the final modality.**

According to the `Lordog/dive-into-llms` repository, modern multi-modal systems like NExT-GPT implement an **encoder-LLM-decoder** architecture that extends traditional language models to handle arbitrary input and output modalities. This design allows the model to receive combinations of text, image, audio, and video, and generate corresponding multi-modal responses through a unified processing pipeline.

## The Encoder-LLM-Decoder Architecture

The fundamental architecture departs from text-only transformers by inserting **modality-specific encoders** before the LLM and **modality-aware decoders** after it. As implemented in [`code/model/anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/anyToImageVideoAudio.py), the system treats the LLM as the central reasoning engine while delegating raw signal processing to specialized components.

### Input Encoding with Modality-Specific Encoders

Raw non-text data enters through dedicated encoders that convert pixels or waveforms into dense vectors. For visual inputs, NExT-GPT utilizes the **ImageBind** encoder, which produces visual token embeddings that capture semantic content across image, video, and audio modalities.

Text inputs follow the standard LLM tokenization pathway using the model's native tokenizer (e.g., Vicuna/LLaMA), producing token IDs that skip the external encoder stage.

### Projection and Fusion into LLM Hidden Space

Once encoded, modality embeddings must align with the LLM's internal dimensionality. The system employs **input projection layers** that map encoder outputs into the LLM's hidden space. As described in the repository's architecture documentation, these projections enable concatenation or interleaving of visual and text tokens into a unified sequence.

The [`anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/anyToImageVideoAudio.py) file orchestrates this fusion, creating a combined token stream where visual embeddings occupy token positions alongside text token IDs, allowing the transformer to process them simultaneously.

## How the LLM Core Processes Multi-Modal Inputs

With inputs converted to a unified representation, the core LLM—typically a Vicuna or LLaMA variant—processes the sequence through its standard transformer blocks. The critical insight is that **the LLM treats projected visual embeddings as special tokens** that can attend to and from text tokens during self-attention.

### Self-Attention Across Text and Visual Tokens

During the forward pass through the LLM's transformer layers, attention heads compute relationships between all tokens in the sequence, regardless of original modality. This mechanism allows the model to ground linguistic concepts in visual features—for example, recognizing that a specific region of an image embedding corresponds to the word "cat" in the text prompt.

The repository emphasizes that the LLM's pre-trained semantic knowledge enables it to interpret these visual tokens without requiring extensive architectural modifications, as the projection layers handle the domain translation.

## Generating Text and Images from the LLM Output

Output generation follows distinct pathways depending on the requested modality. The system dynamically routes the LLM's hidden states to either the standard language modeling head or to specialized diffusion decoders.

### Text Generation via Standard Language Head

For text outputs, the LLM applies its standard **language modeling head** to the final hidden states, producing probability distributions over the vocabulary. Token IDs are sampled autoregressively and detokenized into readable text, following the exact same procedure as text-only LLMs.

### Image Generation via Diffusion Decoders

Image generation relies on external diffusion models that the LLM controls through conditioning signals. As implemented in [`code/model/custom_sd.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/custom_sd.py), the system wraps **Stable Diffusion** to synthesize images from the LLM's outputs.

The LLM generates either:
- A **textual description** that serves as the prompt for the diffusion model
- A **special control token** or learned latent representation that directly conditions the denoising process

The [`custom_sd.py`](https://github.com/Lordog/dive-into-llms/blob/main/custom_sd.py) adapter handles the translation from LLM hidden states to diffusion model conditioning vectors, executing the denoising loop to produce pixel-level outputs. Similar adapters exist for audio ([`custom_ad.py`](https://github.com/Lordog/dive-into-llms/blob/main/custom_ad.py) using AudioLDM) and video ([`custom_vd.py`](https://github.com/Lordog/dive-into-llms/blob/main/custom_vd.py) using ZeroScope).

## Training Pipeline for Multi-Modal Alignment

The repository describes a **three-stage training protocol** that gradually aligns the encoders and decoders with the frozen LLM core:

1. **Input Projection Alignment** – Train the input projection layers to translate encoder embeddings into the LLM's embedding space while keeping the LLM frozen.
2. **Output Projection Alignment** – Train the output adapters that map LLM hidden states to diffusion model conditioning spaces.
3. **Instruction Tuning** – Apply **LoRA** (Low-Rank Adaptation) to fine-tune the LLM itself on multi-modal instruction-following datasets, enabling the model to learn when to generate text versus images.

This staged approach prevents catastrophic forgetting of the LLM's linguistic capabilities while teaching it to control external generative models.

## Code Implementation in NExT-GPT

The `Lordog/dive-into-llms` repository provides concrete implementations demonstrating the encoder-LLM-decoder pattern. Below are minimal working examples derived from the codebase.

### Loading a Pre-Trained NExT-GPT Checkpoint

The main model class in [`code/model/anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/anyToImageVideoAudio.py) initializes the complete pipeline:

```python

# demo_app.py (referenced in the repo)

from code.model.anyToImageVideoAudio import MultiModalModel
import torch

model = MultiModalModel.from_pretrained(
    llm_ckpt="ckpt/pretrained_ckpt/vicuna_ckpt/7b_v0",
    imagebind_ckpt="ckpt/pretrained_ckpt/imagebind_ckpt/huge/imagebind_huge.pth",
    delta_ckpt="ckpt/delta_ckpt/nextgpt/7b_tiva_v0"
)
model.eval()

```

### Encoding an Image and Generating a Caption

This example from [`code/inference.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/inference.py) demonstrates the multi-modal round-trip:

```python
from PIL import Image
import torchvision.transforms as T

img = Image.open("samples/cat.jpg")
img_tensor = T.ToTensor()(img).unsqueeze(0)           # [1,3,H,W]

prompt = "Describe this picture in one sentence."
output = model.infer(
    text=prompt,
    image=img_tensor,
    modalities=["text", "image"]                      # request both text and image output

)

print("Caption:", output["text"])
output["image"].save("generated_illustration.png")

```

### Text-to-Image Generation Pipeline

When generating images without visual input, the model creates dummy embeddings and routes the LLM output to the Stable Diffusion adapter in [`code/model/custom_sd.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/custom_sd.py):

```python
prompt = "A futuristic city skyline at sunset, painted in watercolor."
output = model.infer(
    text=prompt,
    modalities=["image"]
)

output["image"].save("city_watercolor.png")

```

The `infer` method internally invokes the projection layers and diffusion decoder, demonstrating how the LLM acts as a controller for external generative models.

## Summary

- **Multi-modal LLMs** employ an **encoder-LLM-decoder** architecture where modality-specific encoders (like ImageBind) convert raw inputs into embeddings that the LLM processes as special tokens.
- **Projection layers** align encoder outputs with the LLM's hidden space, enabling the transformer to perform cross-modal reasoning through standard self-attention mechanisms.
- **Output generation** branches based on modality: text uses the standard language modeling head, while images route through diffusion decoders like Stable Diffusion ([`custom_sd.py`](https://github.com/Lordog/dive-into-llms/blob/main/custom_sd.py)), controlled by LLM-generated conditioning signals.
- **Training occurs in three stages**: input projection alignment, output projection alignment, and LoRA-based instruction tuning, preserving the LLM's linguistic capabilities while teaching multi-modal control.
- The **NExT-GPT** implementation in `Lordog/dive-into-llms` demonstrates this pipeline through [`anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/anyToImageVideoAudio.py), which orchestrates encoders, the Vicuna LLM, and modality-specific diffusion adapters.

## Frequently Asked Questions

### What is the difference between LLM-as-Scheduler and LLM-as-Joint-Part architectures?

**LLM-as-Scheduler** architectures feed the LLM only text tokens; the model outputs textual commands that downstream modules interpret without direct multimodal fusion. **LLM-as-Joint-Part** architectures, like NExT-GPT, use an encoder-LLM-decoder loop where the LLM directly consumes projected multimodal embeddings and generates control signals for diffusion models, enabling end-to-end gradient flow and tighter integration.

### How does NExT-GPT handle image generation specifically?

NExT-GPT routes the LLM's hidden states through [`code/model/custom_sd.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/custom_sd.py), which wraps Stable Diffusion. The LLM either generates a textual description that serves as the diffusion prompt or emits a learned latent token that directly conditions the denoising process. The adapter projects LLM outputs into the diffusion model's conditioning space, executing the denoising loop to produce pixel-level images.

### Can the model process multiple input modalities simultaneously?

Yes. The [`anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/anyToImageVideoAudio.py) implementation concatenates projected embeddings from multiple encoders (ImageBind for images/audio/video, plus text tokenization) into a single token sequence. The LLM's self-attention mechanism processes these interleaved modalities jointly, allowing the model to reason about relationships between text, images, and other modalities within the same context window.

### What training strategy prevents the LLM from forgetting language capabilities during multi-modal adaptation?

The repository employs a **three-stage training protocol** with parameter-efficient fine-tuning. Stages 1 and 2 train only the input and output projection layers while keeping the LLM frozen. Stage 3 applies **LoRA** (Low-Rank Adaptation) to fine-tune the LLM on instruction-following tasks, adjusting only small adapter matrices rather than full weights. This approach preserves pre-trained linguistic knowledge while teaching the model to coordinate with external encoders and decoders.