How to Choose and Configure Optimal Embedding Models for RAG Implementations in MetaGPT
MetaGPT configures embedding models for Retrieval-Augmented Generation through a centralized factory pattern that reads global configuration settings and instantiates LlamaIndex embedding classes based on provider-specific parameters.
Choosing and configuring optimal embedding models for RAG implementations requires understanding how MetaGPT abstracts provider complexity through its configuration system and factory architecture. The framework delegates embedding creation to the RAGEmbeddingFactory class, which dynamically resolves provider types from config2.yaml and returns ready-to-use LlamaIndex embedding instances. This design allows seamless switching between OpenAI, Azure, Gemini, and local Ollama models without code changes.
Understanding MetaGPT's Embedding Architecture
MetaGPT implements a three-layer architecture for embedding management that separates configuration, validation, and instantiation concerns.
Core Configuration Components
The system relies on three primary files to manage embedding settings:
metagpt/config2.py– Defines the globalConfigclass that loads and exposes settings fromconfig2.yaml, including theembeddingsection.metagpt/configs/embedding_config.py– Contains theEmbeddingConfigPydantic model that validates fields likeapi_type,model,embed_batch_size, anddimensions.metagpt/rag/factories/embedding.py– Houses theRAGEmbeddingFactoryclass and theget_rag_embedding()helper function used throughout the codebase.
The Factory Pattern Implementation
The RAGEmbeddingFactory resolves embedding providers through a four-step process:
-
Provider Resolution – The
_resolve_embedding_type()method checksconfig.embedding.api_typefirst, falling back to legacyconfig.llm.api_typefor backward compatibility. If neither is set, it raises aTypeErrorprompting configuration of an embedding type. -
Creator Selection – Based on the resolved
EmbeddingType, the factory selects a specific creator method:EmbeddingType.OPENAI→_create_openai(usesOpenAIEmbedding)EmbeddingType.AZURE→_create_azure(usesAzureOpenAIEmbedding)EmbeddingType.GEMINI→_create_gemini(usesGeminiEmbedding)EmbeddingType.OLLAMA→_create_ollama(usesOllamaEmbedding)
-
Parameter Construction – Each creator builds a
paramsdictionary from configuration values:params = { "api_key": self.config.embedding.api_key or self.config.llm.api_key, "api_base": self.config.embedding.base_url or self.config.llm.base_url, }The helper
_try_set_model_and_batch_sizeconditionally addsmodel_nameandembed_batch_sizeonly when explicitly configured, preventing unintentional override of provider defaults. -
Instantiation – The factory returns a
BaseEmbeddinginstance by calling the LlamaIndex class with**params.
Configuring Embedding Models in config2.yaml
All embedding configuration resides in the embedding block of config2.yaml (located at metagpt/config/config2.yaml or ~/.metagpt/config2.yaml).
Required Parameters
At minimum, you must specify the provider type:
embedding:
api_type: openai # Options: openai, azure, gemini, ollama
Optional Optimization Parameters
For production RAG implementations, configure these additional fields:
model– Specific model name (e.g.,text-embedding-3-large,text-embedding-ada-002)embed_batch_size– Integer value for bulk indexing operations (e.g.,64,128)dimensions– Vector dimensionality matching your chosen model (e.g.,1536,3072)base_url– Custom endpoint URL for Azure or Ollama deployments
Example configuration for high-dimensional OpenAI embeddings:
embedding:
api_type: openai
api_key: ${OPENAI_API_KEY} # Use environment variables for security
model: text-embedding-3-large
embed_batch_size: 128
dimensions: 3072
Supported Embedding Providers
MetaGPT supports four primary embedding providers, each with specific configuration requirements.
OpenAI
The most mature option with broad model support. Configure in config2.yaml:
embedding:
api_type: openai
model: text-embedding-3-large # or text-embedding-ada-002
embed_batch_size: 64
The factory's _create_openai method (lines 65-75 in metagpt/rag/factories/embedding.py) constructs OpenAIEmbedding instances with automatic API key resolution from either the embedding or LLM configuration blocks.
Azure OpenAI
Enterprise-grade option with VNet compatibility:
embedding:
api_type: azure
api_key: ${AZURE_OPENAI_API_KEY}
base_url: https://my-resource.openai.azure.com
model: text-embedding-3-large
The _create_azure method handles Azure-specific authentication and endpoint configuration.
Google Gemini
Google-centric implementations:
embedding:
api_type: gemini
api_key: ${GEMINI_API_KEY}
model: models/embedding-001
Ollama (Local Deployment)
Cost-effective local execution without API keys:
embedding:
api_type: ollama
base_url: http://localhost:11434
model: nomic-embed-text
embed_batch_size: 32
The _create_ollama method (lines 87-95) specifically handles local endpoint configuration and model naming conventions unique to Ollama deployments.
Implementation Examples
Basic Configuration and Retrieval
Load settings and obtain an embedding instance:
from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding
# Reload configuration to pick up YAML changes
cfg = Config.default(reload=True)
# Get configured embedding model
embedding = get_rag_embedding(config=cfg)
# Generate embedding
vector = await embedding.aget_text_embedding("MetaGPT enables multi-agent collaboration")
print(f"Vector dimensions: {len(vector)}") # Output: 3072 (for text-embedding-3-large)
The get_rag_embedding() helper (defined at lines 110-112 in metagpt/rag/factories/embedding.py) provides a one-liner used by SimpleEngine and other RAG components.
RAG Pipeline Integration
Use the configured embedding within a complete retrieval pipeline:
from metagpt.rag.engines import SimpleEngine
from metagpt.rag.factories import get_rag_embedding, FAISSRetrieverConfig
from llama_index.core.node_parser import SentenceSplitter
# Obtain embedding from factory
embed_model = get_rag_embedding()
# Build retrieval engine
engine = SimpleEngine.from_docs(
input_files=["data/knowledge_base.txt"],
retriever_configs=[FAISSRetrieverConfig()],
transformations=[SentenceSplitter(chunk_size=1024, chunk_overlap=0)],
)
# Execute retrieval
query = "What are the system architecture requirements?"
nodes = await engine.aretrieve(query)
The SimpleEngine.from_docs() method (in metagpt/rag/engines/simple.py, line 383) automatically calls get_rag_embedding() when no explicit embed_model parameter is provided.
Runtime Provider Switching
Dynamically change providers without modifying configuration files:
from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding
cfg = Config.default()
cfg.embedding.api_type = "ollama"
cfg.embedding.base_url = "http://localhost:11434"
cfg.embedding.model = "mxbai-embed-large"
cfg.embedding.embed_batch_size = 16
# Instantiate with new settings
local_embedding = get_rag_embedding(config=cfg)
Performance Optimization Strategies
Batch Size Tuning
Increase embed_batch_size for bulk indexing operations to reduce API round-trips. However, respect provider limits—OpenAI typically handles 2048 items per batch, while local Ollama instances may require smaller values (16-32) depending on hardware constraints.
Dimensionality Selection
Higher-dimensional models (text-embedding-3-large at 3072 dimensions) provide richer semantic representations but increase storage costs and retrieval latency. For FAISS indexes used by SimpleEngine, ensure the dimensions configuration matches your vector store's expected input size.
Cost vs. Accuracy Benchmarking
Evaluate embedding quality using the benchmark script at examples/rag/rag_bm.py. This script uses get_rag_embedding() to test different providers against recall and semantic similarity metrics, allowing data-driven selection based on your quality thresholds and budget constraints.
Summary
- MetaGPT uses a factory pattern (
RAGEmbeddingFactoryinmetagpt/rag/factories/embedding.py) to abstract embedding provider complexity and return standardized LlamaIndexBaseEmbeddinginstances. - Configuration is centralized in
config2.yamlthrough theEmbeddingConfigPydantic model, supporting OpenAI, Azure, Gemini, and Ollama providers. - The
get_rag_embedding()helper provides a singleton-style accessor used bySimpleEngineand other RAG components throughout the framework. - Optimization parameters include
embed_batch_sizefor throughput tuning anddimensionsfor vector store alignment. - Backward compatibility is maintained through fallback to
llm.api_type, though new implementations should explicitly setembedding.api_type.
Frequently Asked Questions
How do I switch from OpenAI to a local Ollama embedding model?
Modify the embedding block in config2.yaml to specify api_type: ollama and provide the local endpoint URL. The RAGEmbeddingFactory automatically selects the _create_ollama method and instantiates OllamaEmbedding from LlamaIndex. Ensure your Ollama server is running and the specified model is pulled locally.
What happens if I don't specify an embedding model name?
If the model field is omitted, the factory's _try_set_model_and_batch_size helper skips adding the parameter to the constructor call. This allows the underlying LlamaIndex embedding class to use its default model (e.g., text-embedding-ada-002 for OpenAI), preventing unintentional overrides while maintaining backward compatibility.
Can I use different embedding models for different RAG pipelines in the same project?
Yes, though the global Config singleton provides one default embedding, you can instantiate multiple configurations programmatically. Create distinct Config instances with different embedding settings and pass them to get_rag_embedding(config=custom_cfg) to obtain provider-specific embedding objects for different pipeline components.
Why does MetaGPT use LlamaIndex embedding classes instead of direct API calls?
MetaGPT delegates to LlamaIndex's BaseEmbedding hierarchy to ensure compatibility with the broader LlamaIndex ecosystem, including vector stores like FAISS and Chroma used by SimpleEngine. This abstraction allows MetaGPT's RAG components to work with any embedding provider that implements the LlamaIndex interface, facilitating seamless provider swaps without engine refactoring.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →