Embedding Models Supported by pathway.xpacks.llm for Vector Indexing
The pathway.xpacks.llm module provides two production-ready embedder implementations—OpenAIEmbedder for cloud-based OpenAI models and SentenceTransformerEmbedder for local HuggingFace sentence-transformers—both fully compatible with Pathway's vector indexing pipeline.
The pathway.xpacks.llm library is part of the Pathway LLM App repository and offers native support for converting text chunks into dense vector representations. Understanding which embedding models are supported by pathway.xpacks.llm for vector indexing is essential for building efficient retrieval-augmented generation (RAG) pipelines and semantic search applications.
Built-in Embedding Models in pathway.xpacks.llm
The package ships with two concrete embedder classes located in the embedders submodule. Both implement a consistent interface that accepts raw text and returns NumPy or PyTorch tensors, allowing seamless integration with pathway.xpacks.llm's Document Store and Vector Store components.
| Embedder Class | Provider | Key Characteristic |
|---|---|---|
OpenAIEmbedder |
OpenAI API | Cloud-based, requires OPENAI_API_KEY |
SentenceTransformerEmbedder |
HuggingFace | Local execution, no API key required |
OpenAIEmbedder: Cloud-Based Embeddings
The OpenAIEmbedder class wraps OpenAI's text embedding models, defaulting to text-embedding-ada-002 while supporting any model available through the OpenAI API.
Configuration and API Requirements
Instantiation requires a valid OpenAI API key set in the environment variable OPENAI_API_KEY. The embedder handles batching and rate limiting automatically when processing text chunks for vector indexing.
In templates/drive_alert/app.py at line 36, the embedder is imported for use in the drive-alert template:
from pathway.xpacks.llm.embedders import OpenAIEmbedder
The YAML configuration in templates/drive_alert/app.yaml at line 59 declares the embedder for the pipeline:
$embedder: !pw.xpacks.llm.embedders.OpenAIEmbedder
Usage Example
import os
from pathway.xpacks.llm.embedders import OpenAIEmbedder
# Ensure your API key is set
os.environ["OPENAI_API_KEY"] = "sk-..."
# Initialize with default ada-002 model
embedder = OpenAIEmbedder()
# Generate embeddings for vector indexing
documents = ["Pathway enables real-time data pipelines.", "Vector indexing supports semantic search."]
embeddings = embedder(documents) # Returns numpy array of shape (2, 1536)
SentenceTransformerEmbedder: Local Open-Source Embeddings
The SentenceTransformerEmbedder class provides fully local embedding generation using HuggingFace's sentence-transformers library, eliminating external API dependencies and latency.
Supported Models and Hardware Options
This embedder accepts any valid HuggingFace model identifier from the sentence-transformers ecosystem, such as all-MiniLM-L6-v2 or avsolatorio/GIST-small-Embedding-v0. You can specify CPU or CUDA device placement via the device parameter.
The templates/private_rag/app.yaml file at line 65 demonstrates configuration for the private RAG template:
$embedder: !pw.xpacks.llm.embedders.SentenceTransformerEmbedder
model: "avsolatorio/GIST-small-Embedding-v0"
Similarly, templates/document_indexing/app.yaml utilizes this embedder for document indexing pipelines.
Usage Example
from pathway.xpacks.llm.embedders import SentenceTransformerEmbedder
# Initialize with a lightweight local model
embedder = SentenceTransformerEmbedder(
model_name="avsolatorio/GIST-small-Embedding-v0",
device="cpu"
)
# Embed documents for vector indexing
chunks = ["Local embeddings ensure data privacy.", "No API calls required."]
vectors = embedder(chunks) # Returns numpy array of shape (2, 384)
The templates/slides_ai_search/README.md at line 311 provides additional guidance on replacing cloud embedders with local sentence-transformer models.
Configuring Embedders via YAML Pipelines
Pathway supports declarative configuration of embedding models through YAML, enabling environment-specific swaps without code changes. The embedder key uses the !pw.xpacks.llm.embedders namespace prefix.
# Configuration for cloud-based embeddings
$embedder: !pw.xpacks.llm.embedders.OpenAIEmbedder
model: "text-embedding-3-small" # Optional: override default ada-002
# Configuration for local embeddings
$embedder: !pw.xpacks.llm.embedders.SentenceTransformerEmbedder
model: "all-MiniLM-L6-v2"
device: "cuda"
Both configurations are compatible with Pathway's vector indexing components, automatically handling batch processing and tensor normalization.
Extending pathway.xpacks.llm with Custom Embedders
While pathway.xpacks.llm officially supports OpenAI and SentenceTransformer embedders, the architecture accepts any callable implementing the embedder protocol. A valid custom embedder must accept a list of strings and return a NumPy array or PyTorch tensor of shape (batch_size, embedding_dimension).
This extensibility allows integration of additional providers such as Cohere, Azure OpenAI, or custom ONNX models without modifying the core library.
Summary
- OpenAIEmbedder provides cloud-based embeddings via OpenAI's API, requiring an
OPENAI_API_KEYand defaulting totext-embedding-ada-002. - SentenceTransformerEmbedder enables fully local, open-source embeddings using any HuggingFace sentence-transformers model without API dependencies.
- Both classes implement the standard embedder interface required by Pathway's vector indexing pipeline, allowing seamless swapping via YAML configuration or Python code.
- The embedders are utilized in production templates including
drive_alert,private_rag, anddocument_indexing.
Frequently Asked Questions
What embedding models are supported by pathway.xpacks.llm for vector indexing?
The pathway.xpacks.llm module officially supports two embedding implementations: OpenAIEmbedder for OpenAI models (such as text-embedding-ada-002 and text-embedding-3-small) and SentenceTransformerEmbedder for any HuggingFace sentence-transformers model (such as all-MiniLM-L6-v2 or avsolatorio/GIST-small-Embedding-v0).
Can I use a custom embedding model not listed in the documentation?
Yes, the embedder interface in pathway.xpacks.llm is designed to be extensible. You can implement a custom embedder by creating a callable class that accepts a list of strings and returns a NumPy array or PyTorch tensor of shape (batch_size, embedding_dimension). This allows integration of providers like Cohere, Azure OpenAI, or proprietary ONNX models.
How do I switch between OpenAI and local embeddings in a Pathway pipeline?
You can switch embedders by modifying the YAML configuration file or the Python instantiation code. In YAML, change the $embedder declaration from !pw.xpacks.llm.embedders.OpenAIEmbedder to !pw.xpacks.llm.embedders.SentenceTransformerEmbedder and specify the desired model name. In Python, simply import and instantiate the alternative class. Both methods maintain compatibility with Pathway's vector indexing components.
Do I need an API key to use pathway.xpacks.llm embedders?
You only need an API key when using OpenAIEmbedder, which requires the OPENAI_API_KEY environment variable to authenticate with OpenAI's cloud service. The SentenceTransformerEmbedder operates entirely locally using HuggingFace models and does not require any external API keys or network access, making it suitable for air-gapped or privacy-sensitive deployments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →