How to Choose and Configure Optimal Embedding Models for RAG Implementations in MetaGPT

MetaGPT configures embedding models for Retrieval-Augmented Generation through a centralized factory pattern that reads global configuration settings and instantiates LlamaIndex embedding classes based on provider-specific parameters.

Choosing and configuring optimal embedding models for RAG implementations requires understanding how MetaGPT abstracts provider complexity through its configuration system and factory architecture. The framework delegates embedding creation to the RAGEmbeddingFactory class, which dynamically resolves provider types from config2.yaml and returns ready-to-use LlamaIndex embedding instances. This design allows seamless switching between OpenAI, Azure, Gemini, and local Ollama models without code changes.

Understanding MetaGPT's Embedding Architecture

MetaGPT implements a three-layer architecture for embedding management that separates configuration, validation, and instantiation concerns.

Core Configuration Components

The system relies on three primary files to manage embedding settings:

The Factory Pattern Implementation

The RAGEmbeddingFactory resolves embedding providers through a four-step process:

  1. Provider Resolution – The _resolve_embedding_type() method checks config.embedding.api_type first, falling back to legacy config.llm.api_type for backward compatibility. If neither is set, it raises a TypeError prompting configuration of an embedding type.

  2. Creator Selection – Based on the resolved EmbeddingType, the factory selects a specific creator method:

    • EmbeddingType.OPENAI → _create_openai (uses OpenAIEmbedding)
    • EmbeddingType.AZURE → _create_azure (uses AzureOpenAIEmbedding)
    • EmbeddingType.GEMINI → _create_gemini (uses GeminiEmbedding)
    • EmbeddingType.OLLAMA → _create_ollama (uses OllamaEmbedding)
  3. Parameter Construction – Each creator builds a params dictionary from configuration values:

    params = {
        "api_key": self.config.embedding.api_key or self.config.llm.api_key,
        "api_base": self.config.embedding.base_url or self.config.llm.base_url,
    }

    The helper _try_set_model_and_batch_size conditionally adds model_name and embed_batch_size only when explicitly configured, preventing unintentional override of provider defaults.

  4. Instantiation – The factory returns a BaseEmbedding instance by calling the LlamaIndex class with **params.

Configuring Embedding Models in config2.yaml

All embedding configuration resides in the embedding block of config2.yaml (located at metagpt/config/config2.yaml or ~/.metagpt/config2.yaml).

Required Parameters

At minimum, you must specify the provider type:

embedding:
  api_type: openai  # Options: openai, azure, gemini, ollama

Optional Optimization Parameters

For production RAG implementations, configure these additional fields:

  • model – Specific model name (e.g., text-embedding-3-large, text-embedding-ada-002)
  • embed_batch_size – Integer value for bulk indexing operations (e.g., 64, 128)
  • dimensions – Vector dimensionality matching your chosen model (e.g., 1536, 3072)
  • base_url – Custom endpoint URL for Azure or Ollama deployments

Example configuration for high-dimensional OpenAI embeddings:

embedding:
  api_type: openai
  api_key: ${OPENAI_API_KEY}  # Use environment variables for security

  model: text-embedding-3-large
  embed_batch_size: 128
  dimensions: 3072

Supported Embedding Providers

MetaGPT supports four primary embedding providers, each with specific configuration requirements.

OpenAI

The most mature option with broad model support. Configure in config2.yaml:

embedding:
  api_type: openai
  model: text-embedding-3-large  # or text-embedding-ada-002

  embed_batch_size: 64

The factory's _create_openai method (lines 65-75 in metagpt/rag/factories/embedding.py) constructs OpenAIEmbedding instances with automatic API key resolution from either the embedding or LLM configuration blocks.

Azure OpenAI

Enterprise-grade option with VNet compatibility:

embedding:
  api_type: azure
  api_key: ${AZURE_OPENAI_API_KEY}
  base_url: https://my-resource.openai.azure.com
  model: text-embedding-3-large

The _create_azure method handles Azure-specific authentication and endpoint configuration.

Google Gemini

Google-centric implementations:

embedding:
  api_type: gemini
  api_key: ${GEMINI_API_KEY}
  model: models/embedding-001

Ollama (Local Deployment)

Cost-effective local execution without API keys:

embedding:
  api_type: ollama
  base_url: http://localhost:11434
  model: nomic-embed-text
  embed_batch_size: 32

The _create_ollama method (lines 87-95) specifically handles local endpoint configuration and model naming conventions unique to Ollama deployments.

Implementation Examples

Basic Configuration and Retrieval

Load settings and obtain an embedding instance:

from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding

# Reload configuration to pick up YAML changes

cfg = Config.default(reload=True)

# Get configured embedding model

embedding = get_rag_embedding(config=cfg)

# Generate embedding

vector = await embedding.aget_text_embedding("MetaGPT enables multi-agent collaboration")
print(f"Vector dimensions: {len(vector)}")  # Output: 3072 (for text-embedding-3-large)

The get_rag_embedding() helper (defined at lines 110-112 in metagpt/rag/factories/embedding.py) provides a one-liner used by SimpleEngine and other RAG components.

RAG Pipeline Integration

Use the configured embedding within a complete retrieval pipeline:

from metagpt.rag.engines import SimpleEngine
from metagpt.rag.factories import get_rag_embedding, FAISSRetrieverConfig
from llama_index.core.node_parser import SentenceSplitter

# Obtain embedding from factory

embed_model = get_rag_embedding()

# Build retrieval engine

engine = SimpleEngine.from_docs(
    input_files=["data/knowledge_base.txt"],
    retriever_configs=[FAISSRetrieverConfig()],
    transformations=[SentenceSplitter(chunk_size=1024, chunk_overlap=0)],
)

# Execute retrieval

query = "What are the system architecture requirements?"
nodes = await engine.aretrieve(query)

The SimpleEngine.from_docs() method (in metagpt/rag/engines/simple.py, line 383) automatically calls get_rag_embedding() when no explicit embed_model parameter is provided.

Runtime Provider Switching

Dynamically change providers without modifying configuration files:

from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding

cfg = Config.default()
cfg.embedding.api_type = "ollama"
cfg.embedding.base_url = "http://localhost:11434"
cfg.embedding.model = "mxbai-embed-large"
cfg.embedding.embed_batch_size = 16

# Instantiate with new settings

local_embedding = get_rag_embedding(config=cfg)

Performance Optimization Strategies

Batch Size Tuning

Increase embed_batch_size for bulk indexing operations to reduce API round-trips. However, respect provider limits—OpenAI typically handles 2048 items per batch, while local Ollama instances may require smaller values (16-32) depending on hardware constraints.

Dimensionality Selection

Higher-dimensional models (text-embedding-3-large at 3072 dimensions) provide richer semantic representations but increase storage costs and retrieval latency. For FAISS indexes used by SimpleEngine, ensure the dimensions configuration matches your vector store's expected input size.

Cost vs. Accuracy Benchmarking

Evaluate embedding quality using the benchmark script at examples/rag/rag_bm.py. This script uses get_rag_embedding() to test different providers against recall and semantic similarity metrics, allowing data-driven selection based on your quality thresholds and budget constraints.

Summary

  • MetaGPT uses a factory pattern (RAGEmbeddingFactory in metagpt/rag/factories/embedding.py) to abstract embedding provider complexity and return standardized LlamaIndex BaseEmbedding instances.
  • Configuration is centralized in config2.yaml through the EmbeddingConfig Pydantic model, supporting OpenAI, Azure, Gemini, and Ollama providers.
  • The get_rag_embedding() helper provides a singleton-style accessor used by SimpleEngine and other RAG components throughout the framework.
  • Optimization parameters include embed_batch_size for throughput tuning and dimensions for vector store alignment.
  • Backward compatibility is maintained through fallback to llm.api_type, though new implementations should explicitly set embedding.api_type.

Frequently Asked Questions

How do I switch from OpenAI to a local Ollama embedding model?

Modify the embedding block in config2.yaml to specify api_type: ollama and provide the local endpoint URL. The RAGEmbeddingFactory automatically selects the _create_ollama method and instantiates OllamaEmbedding from LlamaIndex. Ensure your Ollama server is running and the specified model is pulled locally.

What happens if I don't specify an embedding model name?

If the model field is omitted, the factory's _try_set_model_and_batch_size helper skips adding the parameter to the constructor call. This allows the underlying LlamaIndex embedding class to use its default model (e.g., text-embedding-ada-002 for OpenAI), preventing unintentional overrides while maintaining backward compatibility.

Can I use different embedding models for different RAG pipelines in the same project?

Yes, though the global Config singleton provides one default embedding, you can instantiate multiple configurations programmatically. Create distinct Config instances with different embedding settings and pass them to get_rag_embedding(config=custom_cfg) to obtain provider-specific embedding objects for different pipeline components.

Why does MetaGPT use LlamaIndex embedding classes instead of direct API calls?

MetaGPT delegates to LlamaIndex's BaseEmbedding hierarchy to ensure compatibility with the broader LlamaIndex ecosystem, including vector stores like FAISS and Chroma used by SimpleEngine. This abstraction allows MetaGPT's RAG components to work with any embedding provider that implements the LlamaIndex interface, facilitating seamless provider swaps without engine refactoring.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →