# How to Choose and Configure Optimal Embedding Models for RAG Implementations in MetaGPT

> Master optimal embedding models for RAG in MetaGPT. Learn how MetaGPT's factory pattern configures LlamaIndex embedding classes for efficient retrieval.

- Repository: [FoundationAgents/MetaGPT](https://github.com/FoundationAgents/MetaGPT)
- Tags: best-practices
- Published: 2026-03-04

---

**MetaGPT configures embedding models for Retrieval-Augmented Generation through a centralized factory pattern that reads global configuration settings and instantiates LlamaIndex embedding classes based on provider-specific parameters.**

Choosing and configuring optimal embedding models for RAG implementations requires understanding how MetaGPT abstracts provider complexity through its configuration system and factory architecture. The framework delegates embedding creation to the `RAGEmbeddingFactory` class, which dynamically resolves provider types from [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) and returns ready-to-use LlamaIndex embedding instances. This design allows seamless switching between OpenAI, Azure, Gemini, and local Ollama models without code changes.

## Understanding MetaGPT's Embedding Architecture

MetaGPT implements a three-layer architecture for embedding management that separates configuration, validation, and instantiation concerns.

### Core Configuration Components

The system relies on three primary files to manage embedding settings:

- **[`metagpt/config2.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/config2.py)** – Defines the global `Config` class that loads and exposes settings from [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml), including the `embedding` section.
- **[`metagpt/configs/embedding_config.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/configs/embedding_config.py)** – Contains the `EmbeddingConfig` Pydantic model that validates fields like `api_type`, `model`, `embed_batch_size`, and `dimensions`.
- **[`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py)** – Houses the `RAGEmbeddingFactory` class and the `get_rag_embedding()` helper function used throughout the codebase.

### The Factory Pattern Implementation

The `RAGEmbeddingFactory` resolves embedding providers through a four-step process:

1. **Provider Resolution** – The `_resolve_embedding_type()` method checks `config.embedding.api_type` first, falling back to legacy `config.llm.api_type` for backward compatibility. If neither is set, it raises a `TypeError` prompting configuration of an embedding type.

2. **Creator Selection** – Based on the resolved `EmbeddingType`, the factory selects a specific creator method:
   - `EmbeddingType.OPENAI` → `_create_openai` (uses `OpenAIEmbedding`)
   - `EmbeddingType.AZURE` → `_create_azure` (uses `AzureOpenAIEmbedding`)
   - `EmbeddingType.GEMINI` → `_create_gemini` (uses `GeminiEmbedding`)
   - `EmbeddingType.OLLAMA` → `_create_ollama` (uses `OllamaEmbedding`)

3. **Parameter Construction** – Each creator builds a `params` dictionary from configuration values:
   ```python
   params = {
       "api_key": self.config.embedding.api_key or self.config.llm.api_key,
       "api_base": self.config.embedding.base_url or self.config.llm.base_url,
   }
   ```

   The helper `_try_set_model_and_batch_size` conditionally adds `model_name` and `embed_batch_size` only when explicitly configured, preventing unintentional override of provider defaults.

4. **Instantiation** – The factory returns a `BaseEmbedding` instance by calling the LlamaIndex class with `**params`.

## Configuring Embedding Models in config2.yaml

All embedding configuration resides in the `embedding` block of [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) (located at [`metagpt/config/config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/config/config2.yaml) or `~/.metagpt/config2.yaml`).

### Required Parameters

At minimum, you must specify the provider type:

```yaml
embedding:
  api_type: openai  # Options: openai, azure, gemini, ollama

```

### Optional Optimization Parameters

For production RAG implementations, configure these additional fields:

- **`model`** – Specific model name (e.g., `text-embedding-3-large`, `text-embedding-ada-002`)
- **`embed_batch_size`** – Integer value for bulk indexing operations (e.g., `64`, `128`)
- **`dimensions`** – Vector dimensionality matching your chosen model (e.g., `1536`, `3072`)
- **`base_url`** – Custom endpoint URL for Azure or Ollama deployments

Example configuration for high-dimensional OpenAI embeddings:

```yaml
embedding:
  api_type: openai
  api_key: ${OPENAI_API_KEY}  # Use environment variables for security

  model: text-embedding-3-large
  embed_batch_size: 128
  dimensions: 3072

```

## Supported Embedding Providers

MetaGPT supports four primary embedding providers, each with specific configuration requirements.

### OpenAI

The most mature option with broad model support. Configure in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml):

```yaml
embedding:
  api_type: openai
  model: text-embedding-3-large  # or text-embedding-ada-002

  embed_batch_size: 64

```

The factory's `_create_openai` method (lines 65-75 in [`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py)) constructs `OpenAIEmbedding` instances with automatic API key resolution from either the embedding or LLM configuration blocks.

### Azure OpenAI

Enterprise-grade option with VNet compatibility:

```yaml
embedding:
  api_type: azure
  api_key: ${AZURE_OPENAI_API_KEY}
  base_url: https://my-resource.openai.azure.com
  model: text-embedding-3-large

```

The `_create_azure` method handles Azure-specific authentication and endpoint configuration.

### Google Gemini

Google-centric implementations:

```yaml
embedding:
  api_type: gemini
  api_key: ${GEMINI_API_KEY}
  model: models/embedding-001

```

### Ollama (Local Deployment)

Cost-effective local execution without API keys:

```yaml
embedding:
  api_type: ollama
  base_url: http://localhost:11434
  model: nomic-embed-text
  embed_batch_size: 32

```

The `_create_ollama` method (lines 87-95) specifically handles local endpoint configuration and model naming conventions unique to Ollama deployments.

## Implementation Examples

### Basic Configuration and Retrieval

Load settings and obtain an embedding instance:

```python
from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding

# Reload configuration to pick up YAML changes

cfg = Config.default(reload=True)

# Get configured embedding model

embedding = get_rag_embedding(config=cfg)

# Generate embedding

vector = await embedding.aget_text_embedding("MetaGPT enables multi-agent collaboration")
print(f"Vector dimensions: {len(vector)}")  # Output: 3072 (for text-embedding-3-large)

```

The `get_rag_embedding()` helper (defined at lines 110-112 in [`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py)) provides a one-liner used by `SimpleEngine` and other RAG components.

### RAG Pipeline Integration

Use the configured embedding within a complete retrieval pipeline:

```python
from metagpt.rag.engines import SimpleEngine
from metagpt.rag.factories import get_rag_embedding, FAISSRetrieverConfig
from llama_index.core.node_parser import SentenceSplitter

# Obtain embedding from factory

embed_model = get_rag_embedding()

# Build retrieval engine

engine = SimpleEngine.from_docs(
    input_files=["data/knowledge_base.txt"],
    retriever_configs=[FAISSRetrieverConfig()],
    transformations=[SentenceSplitter(chunk_size=1024, chunk_overlap=0)],
)

# Execute retrieval

query = "What are the system architecture requirements?"
nodes = await engine.aretrieve(query)

```

The `SimpleEngine.from_docs()` method (in [`metagpt/rag/engines/simple.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/engines/simple.py), line 383) automatically calls `get_rag_embedding()` when no explicit `embed_model` parameter is provided.

### Runtime Provider Switching

Dynamically change providers without modifying configuration files:

```python
from metagpt.config2 import Config
from metagpt.rag.factories import get_rag_embedding

cfg = Config.default()
cfg.embedding.api_type = "ollama"
cfg.embedding.base_url = "http://localhost:11434"
cfg.embedding.model = "mxbai-embed-large"
cfg.embedding.embed_batch_size = 16

# Instantiate with new settings

local_embedding = get_rag_embedding(config=cfg)

```

## Performance Optimization Strategies

### Batch Size Tuning

Increase `embed_batch_size` for bulk indexing operations to reduce API round-trips. However, respect provider limits—OpenAI typically handles 2048 items per batch, while local Ollama instances may require smaller values (16-32) depending on hardware constraints.

### Dimensionality Selection

Higher-dimensional models (`text-embedding-3-large` at 3072 dimensions) provide richer semantic representations but increase storage costs and retrieval latency. For FAISS indexes used by `SimpleEngine`, ensure the `dimensions` configuration matches your vector store's expected input size.

### Cost vs. Accuracy Benchmarking

Evaluate embedding quality using the benchmark script at [`examples/rag/rag_bm.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/examples/rag/rag_bm.py). This script uses `get_rag_embedding()` to test different providers against recall and semantic similarity metrics, allowing data-driven selection based on your quality thresholds and budget constraints.

## Summary

- **MetaGPT uses a factory pattern** (`RAGEmbeddingFactory` in [`metagpt/rag/factories/embedding.py`](https://github.com/FoundationAgents/MetaGPT/blob/main/metagpt/rag/factories/embedding.py)) to abstract embedding provider complexity and return standardized LlamaIndex `BaseEmbedding` instances.
- **Configuration is centralized** in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) through the `EmbeddingConfig` Pydantic model, supporting OpenAI, Azure, Gemini, and Ollama providers.
- **The `get_rag_embedding()` helper** provides a singleton-style accessor used by `SimpleEngine` and other RAG components throughout the framework.
- **Optimization parameters** include `embed_batch_size` for throughput tuning and `dimensions` for vector store alignment.
- **Backward compatibility** is maintained through fallback to `llm.api_type`, though new implementations should explicitly set `embedding.api_type`.

## Frequently Asked Questions

### How do I switch from OpenAI to a local Ollama embedding model?

Modify the `embedding` block in [`config2.yaml`](https://github.com/FoundationAgents/MetaGPT/blob/main/config2.yaml) to specify `api_type: ollama` and provide the local endpoint URL. The `RAGEmbeddingFactory` automatically selects the `_create_ollama` method and instantiates `OllamaEmbedding` from LlamaIndex. Ensure your Ollama server is running and the specified model is pulled locally.

### What happens if I don't specify an embedding model name?

If the `model` field is omitted, the factory's `_try_set_model_and_batch_size` helper skips adding the parameter to the constructor call. This allows the underlying LlamaIndex embedding class to use its default model (e.g., `text-embedding-ada-002` for OpenAI), preventing unintentional overrides while maintaining backward compatibility.

### Can I use different embedding models for different RAG pipelines in the same project?

Yes, though the global `Config` singleton provides one default embedding, you can instantiate multiple configurations programmatically. Create distinct `Config` instances with different `embedding` settings and pass them to `get_rag_embedding(config=custom_cfg)` to obtain provider-specific embedding objects for different pipeline components.

### Why does MetaGPT use LlamaIndex embedding classes instead of direct API calls?

MetaGPT delegates to LlamaIndex's `BaseEmbedding` hierarchy to ensure compatibility with the broader LlamaIndex ecosystem, including vector stores like FAISS and Chroma used by `SimpleEngine`. This abstraction allows MetaGPT's RAG components to work with any embedding provider that implements the LlamaIndex interface, facilitating seamless provider swaps without engine refactoring.