Using Groq Universal Sentence Encoder for RAG Evaluation: Implementation Guide
You can evaluate RAG pipelines using Groq's Universal Sentence Encoder by configuring the GroqEmbeddings class from langchain_community and passing it to the evaluate_rag function in evaluation/evalute_rag.py.
The NirDiamant/RAG_Techniques repository provides a modular framework for testing retrieval-augmented generation strategies. This guide demonstrates how to integrate Groq's Universal Sentence Encoder (USE) embedding model and Groq LLM providers into the existing evaluation pipeline defined in the repository's helper functions and evaluation modules.
Understanding the Groq Provider Architecture
The repository abstracts model providers through enums defined in helper_functions.py. This design allows you to swap embedding and LLM backends without modifying the core evaluation logic.
Provider Configuration in helper_functions.py
The ModelProvider enum already includes Groq as a supported LLM provider:
class ModelProvider(Enum):
OPENAI = "openai"
GROQ = "groq"
ANTHROPIC = "anthropic"
AMAZON_BEDROCK = "bedrock"
This definition appears at lines 31-35 in helper_functions.py, enabling you to specify ModelProvider.GROQ when building language model chains.
Setting Up Groq Embeddings for RAG Evaluation
While the repository's get_langchain_embedding_provider factory currently supports OpenAI, Cohere, and Bedrock, extending it for Groq requires minimal modification.
Extending the Embedding Provider Factory
Add a Groq branch to the get_langchain_embedding_provider function in helper_functions.py around line 38:
elif provider == EmbeddingProvider.GROQ:
from langchain_community.embeddings import GroqEmbeddings
return GroqEmbeddings(
model="sentence-transformers/all-MiniLM-L6-v2",
api_key=os.getenv("GROQ_API_KEY")
)
This configuration uses Groq's hosted Universal Sentence Encoder model (all-MiniLM-L6-v2) to generate dense vector representations of your document chunks.
Building the Evaluation Pipeline
With Groq configured as both your embedding and LLM provider, you can execute the full RAG evaluation workflow defined in evaluation/evalute_rag.py.
Vector Store Construction with FAISS
Create a FAISS index using your Groq embeddings:
from langchain_community.vectorstores import FAISS
vector_store = FAISS.from_texts(
documents,
groq_embeddings
)
FAISS provides efficient similarity search for retrieving the most relevant chunks during evaluation.
Configuring the Groq LLM Chain
Initialize the language model and build the question-answering chain:
from langchain_community.chat_models import ChatGroq
from helper_functions import create_question_answer_from_context_chain
groq_llm = ChatGroq(
model="mixtral-8x7b-32768",
api_key="YOUR_GROQ_API_KEY"
)
qa_chain = create_question_answer_from_context_chain(groq_llm)
The create_question_answer_from_context_chain helper (from helper_functions.py) constructs a standard RAG prompt template that injects retrieved context into the query.
Running the Evaluation
Execute the evaluation routine with your configured components:
from evaluation.evalute_rag import evaluate_rag
results = evaluate_rag(
vector_store=vector_store,
qa_chain=qa_chain,
questions=test_questions,
ground_truth=reference_answers,
metrics=["rouge", "bleu", "faithfulness"]
)
The evaluate_rag function retrieves chunks for each question, generates answers using your Groq LLM, and computes metrics defined in evaluation/define_evaluation_metrics.ipynb.
Complete Implementation Example
This self-contained script demonstrates the entire integration:
import os
from helper_functions import (
ModelProvider,
EmbeddingProvider,
create_question_answer_from_context_chain
)
from langchain_community.chat_models import ChatGroq
from langchain_community.embeddings import GroqEmbeddings
from langchain_community.vectorstores import FAISS
from evaluation.evalute_rag import evaluate_rag
# Configuration
documents = ["Your document text here...", "Another chunk..."]
test_questions = ["What is the main topic?"]
ground_truth = ["The main topic is..."]
# 1. Initialize Groq Embeddings (USE)
groq_embeddings = GroqEmbeddings(
model="sentence-transformers/all-MiniLM-L6-v2",
api_key=os.getenv("GROQ_API_KEY")
)
# 2. Build vector store
vector_store = FAISS.from_texts(documents, groq_embeddings)
# 3. Initialize Groq LLM
groq_llm = ChatGroq(
model="mixtral-8x7b-32768",
api_key=os.getenv("GROQ_API_KEY")
)
# 4. Create QA chain
qa_chain = create_question_answer_from_context_chain(groq_llm)
# 5. Run evaluation
results = evaluate_rag(
vector_store=vector_store,
qa_chain=qa_chain,
questions=test_questions,
ground_truth=ground_truth,
metrics=["rouge", "bleu"]
)
print(results)
Summary
- Groq Integration: The repository supports Groq as an LLM provider through
ModelProvider.GROQinhelper_functions.py, enabling sub-second inference for evaluation loops. - Embedding Extension: While the default
get_langchain_embedding_providerrequires a small extension to supportGroqEmbeddings, this unlocks Groq's hosted Universal Sentence Encoder for dense retrieval. - Evaluation Workflow: The
evaluate_ragfunction inevaluation/evalute_rag.pyorchestrates retrieval, generation, and metric computation, accepting any LangChain-compatible vector store and QA chain. - Metrics: Define custom evaluation criteria in
evaluation/define_evaluation_metrics.ipynband pass them as parameters to measure faithfulness, ROUGE, BLEU, or domain-specific accuracy.
Frequently Asked Questions
What is Groq's Universal Sentence Encoder and how does it differ from OpenAI embeddings?
Groq's Universal Sentence Encoder (USE) implementation provides sentence-level embeddings optimized for semantic similarity tasks. Unlike OpenAI's text-embedding-ada-002 which uses a proprietary architecture, Groq's hosted USE model (sentence-transformers/all-MiniLM-L6-v2) is based on the MiniLM architecture, offering competitive performance with lower latency and cost structure suitable for high-throughput RAG evaluation pipelines.
How do I extend the embedding provider factory to support Groq in helper_functions.py?
Locate the get_langchain_embedding_provider function in helper_functions.py around line 38. Add an elif branch checking for EmbeddingProvider.GROQ, then import and return GroqEmbeddings from langchain_community.embeddings. Ensure you pass the model name (e.g., sentence-transformers/all-MiniLM-L6-v2) and API key from environment variables to maintain security best practices.
What evaluation metrics are available in the RAG evaluation pipeline?
The pipeline supports any metric defined in evaluation/define_evaluation_metrics.ipynb, including standard NLP metrics like ROUGE and BLEU for generation quality, as well as RAG-specific metrics such as faithfulness (factual consistency with retrieved context) and answer relevance. You can extend the notebook to implement custom metrics for domain-specific accuracy or latency benchmarks, then pass these as string identifiers to the evaluate_rag function.
Can I use Groq LLM providers with other embedding models in this framework?
Yes, the architecture decouples embedding and LLM providers. You can initialize GroqEmbeddings for dense retrieval while using a different LLM provider (OpenAI, Anthropic, or local models) for answer generation, or conversely use OpenAI embeddings with ChatGroq for inference. The evaluate_rag function in evaluation/evalute_rag.py accepts any LangChain-compatible vector_store and qa_chain regardless of the underlying provider combination.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →