Setting Up Retrieval Augmented Generation (RAG) with Custom Embedding Models in MetaGPT
MetaGPT enables zero-code embedding swaps by configuring the embedding block in config/config2.yaml, which the RAGEmbeddingFactory automatically resolves when SimpleEngine initializes.
MetaGPT ships with a flexible, LlamaIndex-compatible RAG subsystem that supports custom embedding providers like Ollama, Azure, and Gemini through pure configuration. By centralizing embedding setup in config2.yaml, you can switch from OpenAI to local models without changing application code, as the architecture abstracts provider-specific plumbing through factory patterns.
How RAGEmbeddingFactory Resolves Custom Embeddings
The RAGEmbeddingFactory class in metagpt/rag/factories/embedding.py acts as the central resolver for all embedding providers. It maintains a registry of creator methods mapped to specific EmbeddingType and LLMType enums, enabling automatic instantiation based on your configuration.
# metagpt/rag/factories/embedding.py
class RAGEmbeddingFactory(GenericFactory):
def __init__(self, config: Optional[Config] = None):
creators = {
EmbeddingType.OPENAI: self._create_openai,
EmbeddingType.AZURE: self._create_azure,
EmbeddingType.GEMINI: self._create_gemini,
EmbeddingType.OLLAMA: self._create_ollama,
# Backward-compatibility with LLM enum
LLMType.OPENAI: self._create_openai,
LLMType.AZURE: self._create_azure,
}
super().__init__(creators)
self.config = config if config else Config.default()
When get_rag_embedding() is invoked, the factory executes _resolve_embedding_type() to determine which provider to use. If config.embedding.api_type is explicitly set, that value takes precedence. If omitted, the factory falls back to the LLM type for backward compatibility. When neither is defined, the factory raises a TypeError with the message: "To use RAG, please set your embedding in config2.yaml."
Each creator method (_create_openai, _create_ollama, etc.) constructs the concrete BaseEmbedding subclass (e.g., OpenAIEmbedding, OllamaEmbedding) using credentials from the embedding block, with automatic fallback to the LLM block for missing keys. The factory injects optional parameters like model and embed_batch_size via _try_set_model_and_batch_size.
SimpleEngine Auto-Selection Mechanism
The SimpleEngine class in metagpt/rag/engines/simple.py automatically retrieves the appropriate embedding model without manual instantiation. When you call SimpleEngine.from_docs(), the engine delegates embedding resolution to the factory through the _resolve_embed_model static method.
# metagpt/rag/engines/simple.py
@staticmethod
def _resolve_embed_model(embed_model: BaseEmbedding = None, configs: list[Any] = None) -> BaseEmbedding:
if configs and all(isinstance(c, NoEmbedding) for c in configs):
return MockEmbedding(embed_dim=1)
return embed_model or get_rag_embedding()
If you pass an explicit embed_model argument to from_docs(), that instance overrides the factory default. Otherwise, the engine calls get_rag_embedding(), which reads your config2.yaml and returns the configured embedding provider. This design ensures that switching from OpenAI to Ollama requires only YAML changes, not code modifications.
Configuring Custom Embedding Models in Practice
Editing config2.yaml
Add an embedding configuration block to config/config2.yaml alongside your existing llm settings. You only need to specify the api_type, base_url, and model name for most providers.
# config/config2.yaml
llm:
api_type: "openai"
model: "gpt-4-turbo"
base_url: "https://api.openai.com/v1"
api_key: "YOUR_API_KEY"
embedding:
api_type: "ollama" # Provider identifier
base_url: "http://localhost:11434" # Local Ollama endpoint
model: "mxbai-embed-large" # Model name on the server
embed_batch_size: 32 # Optional batch size for inference
The embedding block accepts the same provider types as the LLM configuration, including "openai", "azure", "gemini", and "ollama". Credentials defined here take precedence over the LLM block, though the factory will fall back to LLM keys if embedding-specific credentials are absent.
Building the RAG Engine
With the YAML configured, instantiate SimpleEngine without importing provider-specific classes. The factory handles the instantiation internally.
# my_rag.py
import asyncio
from metagpt.rag.engines import SimpleEngine
from metagpt.rag.schema import FAISSRetrieverConfig
async def run():
# Automatically uses the embedding defined in config2.yaml
engine = SimpleEngine.from_docs(
input_files=["data/my_corpus.txt"],
retriever_configs=[FAISSRetrieverConfig()],
)
# Standard retrieval and query workflow
nodes = await engine.aretrieve("What is the key idea?")
answer = await engine.aquery("What is the key idea?")
print(f"Retrieved {len(nodes)} nodes")
print("Answer:", answer)
if __name__ == "__main__":
asyncio.run(run())
Verifying the Embedding Configuration
To confirm which embedding class is active, call the factory directly before building the engine.
from metagpt.rag.factories import get_rag_embedding
embed = get_rag_embedding() # Resolves from config2.yaml
print(type(embed)) # <class 'llama_index.embeddings.ollama.OllamaEmbedding'>
Complete Working Example
This end-to-end example demonstrates loading documents, building a FAISS index, and querying with a custom Ollama embedding model. The code assumes you have configured config2.yaml as shown in the previous section.
# example_custom_rag.py
import asyncio
from pathlib import Path
from metagpt.rag.engines import SimpleEngine
from metagpt.rag.schema import FAISSRetrieverConfig, LLMRankerConfig
DATA_DIR = Path(__file__).parent / "data"
DOC_PATH = DATA_DIR / "faq.txt"
QUESTION = "How does MetaGPT handle multi-agent collaboration?"
async def main():
# Engine automatically picks the embedding from config2.yaml
engine = SimpleEngine.from_docs(
input_files=[DOC_PATH],
retriever_configs=[FAISSRetrieverConfig()],
ranker_configs=[LLMRankerConfig()], # Optional reranker
)
# Retrieve relevant chunks
chunks = await engine.aretrieve(QUESTION)
print(f"Top {len(chunks)} retrieved chunks:")
for i, node in enumerate(chunks):
print(f"{i}. {node.text[:80]}… (score={node.score:.2f})")
# Generate final answer
answer = await engine.aquery(QUESTION)
print("\n=== Answer ===")
print(answer)
if __name__ == "__main__":
asyncio.run(main())
Place this script adjacent to your document directory, ensure config2.yaml points to your desired embedding endpoint, and execute. The output displays retrieval scores from the custom embedding model followed by the LLM-generated answer.
Summary
-
Declare your embedding provider in
config/config2.yamlusing theembedding.api_typefield; supported values include"openai","azure","gemini", and"ollama". -
RAGEmbeddingFactory resolves the configuration, constructs the appropriate
BaseEmbeddingsubclass, and handles credential fallback between the embedding and LLM configuration blocks. -
SimpleEngine automatically invokes
get_rag_embedding()during initialization, enabling zero-code provider swaps while maintaining the samefrom_docs(),aretrieve(), andaquery()API. -
Override the default embedding per-engine by passing an explicit
embed_modelparameter toSimpleEngine.from_docs()when specific instances require different vectorization logic.
Frequently Asked Questions
What embedding providers does MetaGPT support?
MetaGPT supports any LlamaIndex-compatible embedding provider including OpenAI, Azure OpenAI, Google Gemini, and Ollama. The RAGEmbeddingFactory in metagpt/rag/factories/embedding.py maintains creator methods for each provider, and you activate them by setting embedding.api_type to the corresponding identifier in config2.yaml.
Can I use different embeddings for different engines in the same project?
Yes. While the factory provides a default embedding based on config2.yaml, you can override this per-engine by passing a specific embed_model instance to SimpleEngine.from_docs(embed_model=your_custom_embedding). This allows one engine to use OpenAI embeddings while another uses local Ollama models within the same runtime.
How do I troubleshoot "To use RAG, please set your embedding in config2.yaml"?
This TypeError indicates that RAGEmbeddingFactory._resolve_embedding_type() found neither config.embedding.api_type nor a valid config.llm.api_type to fall back on. Verify that your config/config2.yaml contains an embedding block with at minimum the api_type field defined, or ensure your llm configuration is properly loaded if you intend to use the fallback behavior.
Does MetaGPT support local embedding models without internet access?
Yes. By configuring embedding.api_type: "ollama" and pointing base_url to a local endpoint (e.g., http://localhost:11434), you can run entirely offline. The OllamaEmbedding class from LlamaIndex handles local inference, and MetaGPT's factory automatically instantiates it when it detects the Ollama configuration in config2.yaml.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →