# Architecture of Multi-Modal Large Language Models: Encoder-LLM-Decoder vs Task Scheduler Patterns

> Explore multi-modal large language model architectures. Understand encoder-LLM-decoder and task scheduler patterns for processing diverse data inputs.

- Repository: [Tongxin Yuan/dive-into-llms](https://github.com/Lordog/dive-into-llms)
- Tags: architecture
- Published: 2026-04-16

---

**Multi-modal large language models (MLLMs) typically employ either an encoder-LLM-decoder pipeline where the language model directly consumes cross-modal embeddings, or a task-scheduler pattern where the LLM orchestrates external modality-specific modules through text commands.**

The `Lordog/dive-into-llms` repository provides a comprehensive implementation of both architectural approaches, demonstrating how modern systems extend traditional large language models to process and generate text, images, audio, and video simultaneously. Understanding these architectural patterns is essential for developers building end-to-end multi-modal applications.

## Two Dominant MLLM Architecture Patterns

According to the documentation in [`documents/chapter8/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter8/README.md), the repository identifies two fundamental paradigms for integrating multiple modalities with large language models.

### LLM as Task Scheduler

In this architectural pattern, the LLM receives text prompts and emits **text-only commands** that drive downstream modality-specific modules. The LLM never directly processes non-text data; instead, it acts as a high-level orchestrator that generates instructions for specialized components to execute.

All communication between the LLM and other system components occurs through textual instructions. For example, the model might output a command like "generate_image('sunset over mountains')" which triggers an external image generation module, but the LLM itself only handles the textual representation of that command.

### LLM as Joint Part of System (Encoder-LLM-Decoder)

This architecture—implemented in [`code/model/anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/anyToImageVideoAudio.py)—places the LLM at the center of an **encoder-LLM-decoder pipeline**. The language model directly consumes embeddings from modality encoders and produces embeddings for modality decoders, enabling native understanding and generation across all supported formats.

The `Lordog/dive-into-llms` implementation specifically uses this pattern in its NExT-GPT demonstration, where the Vicuna-based LLM core processes continuous vector representations from ImageBind encoders and outputs latent embeddings that feed diffusion-based decoders.

## The Encoder-LLM-Decoder Pipeline in Detail

The repository's implementation follows a three-stage processing flow that enables *any-to-any* modality conversion (e.g., text + image → audio, or text + video → text + image + audio).

### Modality Encoders

The system utilizes **ImageBind** encoders to process images, audio, and video into a shared embedding space. Located in the `code/model/ImageBind` directory, these encoders convert raw sensory inputs into continuous vector representations that the LLM can process.

Unlike traditional systems that require separate encoders for each modality, ImageBind provides a unified encoding mechanism that maps all modalities into a common semantic space, allowing the LLM to reason across different input types simultaneously.

### LLM Core

At the heart of the architecture sits a **Vicuna-style LLM** built on the LLaMA architecture, implemented in [`code/model/modeling_llama.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/modeling_llama.py). This component receives both the textual instructions and the embedded representations from the modality encoders.

The LLM performs token-level reasoning over the combined inputs, using its attention mechanisms to relate textual tokens with continuous visual and audio embeddings. Based on the input context and instruction, it determines the appropriate output modality and generates corresponding embeddings for the downstream decoders.

### Modality Decoders

The output stage employs specialized diffusion models to generate final content:

- **Stable Diffusion** for image generation
- **AudioLDM** for audio synthesis  
- **ZeroScope** for video generation

These decoders, referenced in [`code/config/base.yaml`](https://github.com/Lordog/dive-into-llms/blob/main/code/config/base.yaml), receive the LLM's output embeddings and convert them back into perceptible media formats. The decoders operate in the same latent space as the encoders, ensuring semantic consistency between what the LLM "understands" and what it generates.

## Implementation: Running NExT-GPT Inference

The repository provides a high-level interface through [`code/demo_app.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/demo_app.py) that demonstrates the encoder-LLM-decoder architecture in action. Below is a minimal Python snippet that loads the model and performs a "text + image → audio" conversion:

```python
from code.demo_app import DemoApp

# Initialise the demo; paths point to the pretrained checkpoints.

app = DemoApp(
    llm_ckpt_path="ckpt/pretrained_ckpt/vicuna_ckpt/7b_v0",
    imagebind_ckpt_path="ckpt/pretrained_ckpt/imagebind_ckpt/huge/imagebind_huge.pth",
    diffusion_ckpt_dir="ckpt/pretrained_ckpt",   # contains Stable Diffusion, AudioLDM, ZeroScope

    delta_ckpt_path="ckpt/delta_ckpt/nextgpt/7b_tiva_v0"
)

# Example input: a textual prompt and an image file.

prompt = "Describe the scene and generate a matching background music."
image_path = "./data/T-X_pair_data/cc3m/images/sample1.jpg"

# Perform inference – the model returns a dict with generated text and audio bytes.

result = app.run(prompt=prompt, image_path=image_path)

print("Generated description:", result["text"])

# `result["audio"]` contains raw WAV bytes; you could write it to a file:

with open("output.wav", "wb") as f:
    f.write(result["audio"])

```

The script first loads frozen ImageBind encoders and the LLM backbone, then applies the learnable *delta* parameters specific to NExT-GPT. During inference, it routes the multimodal inputs through the encoder-LLM-decoder pipeline, returning both textual descriptions and raw audio bytes. You can launch the full Gradio interface by executing `bash scripts/app.sh` from the repository root.

## Key Source Files and Components

The following files embody the architecture of multi-modal large language models as implemented in this repository:

- [`documents/chapter8/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter8/README.md) – Documents the two architectural patterns with visual diagrams
- [`code/model/anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/anyToImageVideoAudio.py) – Core model class wiring encoders, LLM, and decoders
- `code/model/ImageBind/` – Unified image/video/audio encoder implementation
- [`code/model/modeling_llama.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/modeling_llama.py) – LLM backbone handling tokenization and generation
- [`code/demo_app.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/demo_app.py) – High-level wrapper for inference and UI formatting
- [`code/config/base.yaml`](https://github.com/Lordog/dive-into-llms/blob/main/code/config/base.yaml) – Hyper-parameters and component paths
- [`code/inference.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/inference.py) – Command-line entry point for headless inference
- [`scripts/train.sh`](https://github.com/Lordog/dive-into-llms/blob/main/scripts/train.sh) – Training pipeline for three-stage alignment and instruction-tuning

## Summary

- **Multi-modal large language models** primarily follow either an encoder-LLM-decoder architecture or a task-scheduler pattern where the LLM orchestrates external modules via text commands.
- The **encoder-LLM-decoder** design implemented in `Lordog/dive-into-llms` uses ImageBind for unified multimodal encoding, Vicuna/LLaMA for central reasoning, and diffusion decoders for content generation.
- This architecture enables **any-to-any modality conversion**, allowing inputs and outputs to mix text, image, audio, and video in single inference passes.
- The `DemoApp` class in [`code/demo_app.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/demo_app.py) provides a practical interface for running multi-modal inference with support for delta checkpoints and frozen pretrained components.

## Frequently Asked Questions

### What is the primary difference between the task-scheduler and encoder-LLM-decoder architectures?

The **task-scheduler** architecture restricts the LLM to processing text only, using natural language commands to trigger external modality-specific tools, while the **encoder-LLM-decoder** architecture allows the LLM to directly consume and generate continuous embeddings from multiple modalities, creating a unified representation space.

### How does the NExT-GPT implementation handle any-to-any modality conversion?

According to the source code in [`code/model/anyToImageVideoAudio.py`](https://github.com/Lordog/dive-into-llms/blob/main/code/model/anyToImageVideoAudio.py), the system uses ImageBind encoders to project all input modalities into a shared embedding space, processes these through the Vicuna LLM core, and routes the outputs to appropriate diffusion decoders (Stable Diffusion for images, AudioLDM for audio, ZeroScope for video), enabling seamless conversion between any input and output modality combinations.

### Where are the model configurations and checkpoint paths defined in the repository?

The default hyper-parameters and component paths are specified in [`code/config/base.yaml`](https://github.com/Lordog/dive-into-llms/blob/main/code/config/base.yaml), which references checkpoint directories for the Vicuna LLM (`ckpt/pretrained_ckpt/vicuna_ckpt/`), ImageBind encoders (`ckpt/pretrained_ckpt/imagebind_ckpt/`), and diffusion decoders, along with delta checkpoint paths for the tuned NExT-GPT parameters.

### What role does the `delta_ckpt` play in the NExT-GPT architecture?

The delta checkpoint (`ckpt/delta_ckpt/nextgpt/7b_tiva_v0`) contains the trainable parameters added to the frozen pretrained LLM, encoders, and decoders during the three-stage alignment process described in [`scripts/train.sh`](https://github.com/Lordog/dive-into-llms/blob/main/scripts/train.sh), enabling the system to learn cross-modal projections without fine-tuning the entire parameter space.