# Setting Up Retrieval Augmented Generation (RAG) with Custom Embedding Models in MetaGPT

> Learn to set up Retrieval Augmented Generation RAG with custom embedding models in MetaGPT. Easily swap embedding models via configyaml for enhanced AI applications.

- Repository: [FoundationAgents/MetaGPT](https://github.com/FoundationAgents/MetaGPT)
- Tags: how-to-guide
- Published: 2026-03-04

---

**MetaGPT enables zero-code embedding swaps by configuring the `embedding` block in [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml), which the `RAGEmbeddingFactory` automatically resolves when `SimpleEngine` initializes.**

MetaGPT ships with a flexible, LlamaIndex-compatible RAG subsystem that supports custom embedding providers like Ollama, Azure, and Gemini through pure configuration. By centralizing embedding setup in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml), you can switch from OpenAI to local models without changing application code, as the architecture abstracts provider-specific plumbing through factory patterns.

## How RAGEmbeddingFactory Resolves Custom Embeddings

The `RAGEmbeddingFactory` class in [`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py) acts as the central resolver for all embedding providers. It maintains a registry of creator methods mapped to specific `EmbeddingType` and `LLMType` enums, enabling automatic instantiation based on your configuration.

```python

# metagpt/rag/factories/embedding.py

class RAGEmbeddingFactory(GenericFactory):
    def __init__(self, config: Optional[Config] = None):
        creators = {
            EmbeddingType.OPENAI: self._create_openai,
            EmbeddingType.AZURE: self._create_azure,
            EmbeddingType.GEMINI: self._create_gemini,
            EmbeddingType.OLLAMA: self._create_ollama,
            # Backward-compatibility with LLM enum

            LLMType.OPENAI: self._create_openai,
            LLMType.AZURE: self._create_azure,
        }
        super().__init__(creators)
        self.config = config if config else Config.default()

```

When `get_rag_embedding()` is invoked, the factory executes `_resolve_embedding_type()` to determine which provider to use. If `config.embedding.api_type` is explicitly set, that value takes precedence. If omitted, the factory falls back to the LLM type for backward compatibility. When neither is defined, the factory raises a `TypeError` with the message: `"To use RAG, please set your embedding in config2.yaml."`

Each creator method (`_create_openai`, `_create_ollama`, etc.) constructs the concrete `BaseEmbedding` subclass (e.g., `OpenAIEmbedding`, `OllamaEmbedding`) using credentials from the **embedding** block, with automatic fallback to the **LLM** block for missing keys. The factory injects optional parameters like `model` and `embed_batch_size` via `_try_set_model_and_batch_size`.

## SimpleEngine Auto-Selection Mechanism

The `SimpleEngine` class in [`metagpt/rag/engines/simple.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/engines/simple.py) automatically retrieves the appropriate embedding model without manual instantiation. When you call `SimpleEngine.from_docs()`, the engine delegates embedding resolution to the factory through the `_resolve_embed_model` static method.

```python

# metagpt/rag/engines/simple.py

@staticmethod
def _resolve_embed_model(embed_model: BaseEmbedding = None, configs: list[Any] = None) -> BaseEmbedding:
    if configs and all(isinstance(c, NoEmbedding) for c in configs):
        return MockEmbedding(embed_dim=1)

    return embed_model or get_rag_embedding()

```

If you pass an explicit `embed_model` argument to `from_docs()`, that instance overrides the factory default. Otherwise, the engine calls `get_rag_embedding()`, which reads your [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) and returns the configured embedding provider. This design ensures that switching from OpenAI to Ollama requires only YAML changes, not code modifications.

## Configuring Custom Embedding Models in Practice

### Editing config2.yaml

Add an `embedding` configuration block to [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml) alongside your existing `llm` settings. You only need to specify the `api_type`, `base_url`, and `model` name for most providers.

```yaml

# config/config2.yaml

llm:
  api_type: "openai"
  model: "gpt-4-turbo"
  base_url: "https://api.openai.com/v1"
  api_key: "YOUR_API_KEY"

embedding:
  api_type: "ollama"                     # Provider identifier

  base_url: "http://localhost:11434"     # Local Ollama endpoint

  model: "mxbai-embed-large"            # Model name on the server

  embed_batch_size: 32                   # Optional batch size for inference

```

The `embedding` block accepts the same provider types as the LLM configuration, including `"openai"`, `"azure"`, `"gemini"`, and `"ollama"`. Credentials defined here take precedence over the LLM block, though the factory will fall back to LLM keys if embedding-specific credentials are absent.

### Building the RAG Engine

With the YAML configured, instantiate `SimpleEngine` without importing provider-specific classes. The factory handles the instantiation internally.

```python

# my_rag.py

import asyncio
from metagpt.rag.engines import SimpleEngine
from metagpt.rag.schema import FAISSRetrieverConfig

async def run():
    # Automatically uses the embedding defined in config2.yaml

    engine = SimpleEngine.from_docs(
        input_files=["data/my_corpus.txt"],
        retriever_configs=[FAISSRetrieverConfig()],
    )
    
    # Standard retrieval and query workflow

    nodes = await engine.aretrieve("What is the key idea?")
    answer = await engine.aquery("What is the key idea?")
    print(f"Retrieved {len(nodes)} nodes")
    print("Answer:", answer)

if __name__ == "__main__":
    asyncio.run(run())

```

### Verifying the Embedding Configuration

To confirm which embedding class is active, call the factory directly before building the engine.

```python
from metagpt.rag.factories import get_rag_embedding

embed = get_rag_embedding()          # Resolves from config2.yaml

print(type(embed))                    # <class 'llama_index.embeddings.ollama.OllamaEmbedding'>

```

## Complete Working Example

This end-to-end example demonstrates loading documents, building a FAISS index, and querying with a custom Ollama embedding model. The code assumes you have configured [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) as shown in the previous section.

```python

# example_custom_rag.py

import asyncio
from pathlib import Path

from metagpt.rag.engines import SimpleEngine
from metagpt.rag.schema import FAISSRetrieverConfig, LLMRankerConfig

DATA_DIR = Path(__file__).parent / "data"
DOC_PATH = DATA_DIR / "faq.txt"
QUESTION = "How does MetaGPT handle multi-agent collaboration?"

async def main():
    # Engine automatically picks the embedding from config2.yaml

    engine = SimpleEngine.from_docs(
        input_files=[DOC_PATH],
        retriever_configs=[FAISSRetrieverConfig()],
        ranker_configs=[LLMRankerConfig()],   # Optional reranker

    )

    # Retrieve relevant chunks

    chunks = await engine.aretrieve(QUESTION)
    print(f"Top {len(chunks)} retrieved chunks:")
    for i, node in enumerate(chunks):
        print(f"{i}. {node.text[:80]}… (score={node.score:.2f})")

    # Generate final answer

    answer = await engine.aquery(QUESTION)
    print("\n=== Answer ===")
    print(answer)

if __name__ == "__main__":
    asyncio.run(main())

```

Place this script adjacent to your document directory, ensure [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) points to your desired embedding endpoint, and execute. The output displays retrieval scores from the custom embedding model followed by the LLM-generated answer.

## Summary

- **Declare** your embedding provider in [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml) using the `embedding.api_type` field; supported values include `"openai"`, `"azure"`, `"gemini"`, and `"ollama"`.

- **RAGEmbeddingFactory** resolves the configuration, constructs the appropriate `BaseEmbedding` subclass, and handles credential fallback between the embedding and LLM configuration blocks.

- **SimpleEngine** automatically invokes `get_rag_embedding()` during initialization, enabling zero-code provider swaps while maintaining the same `from_docs()`, `aretrieve()`, and `aquery()` API.

- **Override** the default embedding per-engine by passing an explicit `embed_model` parameter to `SimpleEngine.from_docs()` when specific instances require different vectorization logic.

## Frequently Asked Questions

### What embedding providers does MetaGPT support?

MetaGPT supports any LlamaIndex-compatible embedding provider including OpenAI, Azure OpenAI, Google Gemini, and Ollama. The `RAGEmbeddingFactory` in [`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py) maintains creator methods for each provider, and you activate them by setting `embedding.api_type` to the corresponding identifier in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml).

### Can I use different embeddings for different engines in the same project?

Yes. While the factory provides a default embedding based on [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml), you can override this per-engine by passing a specific `embed_model` instance to `SimpleEngine.from_docs(embed_model=your_custom_embedding)`. This allows one engine to use OpenAI embeddings while another uses local Ollama models within the same runtime.

### How do I troubleshoot "To use RAG, please set your embedding in config2.yaml"?

This `TypeError` indicates that `RAGEmbeddingFactory._resolve_embedding_type()` found neither `config.embedding.api_type` nor a valid `config.llm.api_type` to fall back on. Verify that your [`config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config/config2.yaml) contains an `embedding` block with at minimum the `api_type` field defined, or ensure your `llm` configuration is properly loaded if you intend to use the fallback behavior.

### Does MetaGPT support local embedding models without internet access?

Yes. By configuring `embedding.api_type: "ollama"` and pointing `base_url` to a local endpoint (e.g., `http://localhost:11434`), you can run entirely offline. The `OllamaEmbedding` class from LlamaIndex handles local inference, and MetaGPT's factory automatically instantiates it when it detects the Ollama configuration in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml).