Retrieval with Feedback Loops in RAG: Implementing Continuous Improvement in Vector Search
Retrieval with feedback loops in RAG creates a closed-loop system that dynamically improves retrieval quality by incorporating explicit user ratings to re-rank results and periodically re-index high-quality feedback into the vector store.
The NirDiamant/RAG_Techniques repository provides a production-ready implementation of retrieval with feedback loops in RAG, demonstrating how explicit human judgments can continuously refine retrieval performance without retraining underlying embedding models. This technique bridges the gap between static vector indices and dynamic user needs by capturing relevance and quality scores, then using them to adjust ranking algorithms and regenerate the knowledge base.
How Retrieval with Feedback Loops Works in RAG
The architecture follows a four-stage pipeline implemented in all_rag_techniques_runnable_scripts/retrieval_with_feedback_loop.py. Each stage builds upon the previous to create a self-improving retrieval system.
Stage 1: Document Ingestion and Vector Indexing
The pipeline begins by converting raw documents into a searchable vector index. The read_pdf_to_string function in helper_functions.py extracts plain text from PDF files, which is then processed by encode_from_string to create a FAISS vector store. The resulting vectorstore object is wrapped by a LangChain retriever via self.retriever = self.vectorstore.as_retriever(), establishing the initial retrieval interface.
Stage 2: Initial Retrieval and Generation
Once indexed, the system processes queries using a ChatOpenAI LLM (specifically GPT-4o) paired with the retriever through RetrievalQA.from_chain_type. When self.qa_chain(query)["result"] executes, the retriever fetches candidate documents, the LLM synthesizes an answer, and the raw response is presented to the user for evaluation.
Stage 3: Explicit Feedback Capture
The feedback loop activates when users provide numeric ratings for relevance and quality (scaled 1-5) along with optional comments. The get_user_feedback function structures this input, while store_feedback serializes the data and appends it to data/feedback_data.json. This persistent storage creates a growing dataset of human judgments tied to specific queries and retrieved documents.
Stage 4: Dynamic Re-ranking and Index Fine-tuning
The system leverages stored feedback through two mechanisms implemented in the RetrievalAugmentedGeneration class:
Relevance Score Adjustment: The adjust_relevance_scores method reads historical feedback and uses the LLM to judge whether past entries are relevant to the current query via a structured relevance_prompt. When feedback applies, the document's metadata['relevance_score'] is scaled by the average relevance rating: doc.metadata['relevance_score'] *= (avg_relevance / 3). The retriever's search_kwargs['k'] is then updated to reflect the adjusted document set.
Index Fine-tuning: The fine_tune_index method identifies high-quality feedback entries (where relevance ≥ 4 and quality ≥ 4) and concatenates them with the original corpus. The combined text is re-encoded using encode_from_string to produce a refreshed vector store that incorporates successful query-response patterns directly into the searchable knowledge base.
Implementing the Feedback Loop: Code Examples
Command-Line Execution
Run the complete pipeline from the terminal to process a PDF and submit feedback:
python all_rag_techniques_runnable_scripts/retrieval_with_feedback_loop.py \
--path data/Understanding_Climate_Change.pdf \
--query "What are the main causes of climate change?" \
--relevance 5 \
--quality 5
The script outputs the generated answer and persists the feedback entry to data/feedback_data.json.
Programmatic Integration
Import the RetrievalAugmentedGeneration class to embed the feedback loop within larger applications:
from all_rag_techniques_runnable_scripts.retrieval_with_feedback_loop import RetrievalAugmentedGeneration
# Initialize with a document
rag = RetrievalAugmentedGeneration("../data/Understanding_Climate_Change.pdf")
# Execute query with feedback
answer = rag.run(
query="How does deforestation affect carbon levels?",
relevance=4,
quality=5
)
print("Answer:", answer)
# Refresh index after accumulating high-quality feedback
new_vs = rag.fine_tune_index(
feedback_data=rag.load_feedback_data(),
texts=rag.content
)
rag.retriever = new_vs.as_retriever()
Analyzing Collected Feedback
Inspect the feedback dataset to monitor system performance:
import json
from pathlib import Path
feedback_path = Path("../data/feedback_data.json")
feedback = [json.loads(line) for line in feedback_path.read_text().splitlines()]
print(f"Collected {len(feedback)} feedback entries")
Key Files and Architecture
| File | Role | Location |
|---|---|---|
all_rag_techniques_runnable_scripts/retrieval_with_feedback_loop.py |
Main RAG pipeline implementing the RetrievalAugmentedGeneration class, feedback handling, relevance adjustment, and index fine-tuning. |
View on GitHub |
helper_functions.py |
Utility functions for PDF loading (read_pdf_to_string), text chunking, FAISS embedding (encode_from_string), BM25 retrieval, and provider factories. |
View on GitHub |
evaluation/evalute_rag.py |
Evaluation harness using deepeval to benchmark RAG output quality; useful for quantifying the impact of feedback-driven improvements. |
View on GitHub |
README.md |
Repository overview and quick-start instructions for running the various RAG techniques. | View on GitHub |
Summary
- Retrieval with feedback loops in RAG enables continuous improvement by capturing explicit user ratings on relevance and quality, then using that data to dynamically adjust retrieval rankings.
- The implementation in
NirDiamant/RAG_Techniquesuses a four-stage pipeline: document ingestion, initial retrieval with GPT-4o, structured feedback capture todata/feedback_data.json, and dynamic re-ranking viaadjust_relevance_scores. - High-quality feedback entries (scores ≥ 4) are periodically incorporated into the vector store through
fine_tune_index, creating a synthetic data augmentation loop that refreshes the FAISS index without retraining the underlying embedding model. - The modular architecture allows integration with existing LangChain applications, supporting both command-line execution and programmatic use via the
RetrievalAugmentedGenerationclass.
Frequently Asked Questions
How does the feedback loop improve retrieval without retraining embeddings?
The system improves retrieval through dynamic score adjustment and index augmentation rather than model retraining. The adjust_relevance_scores method re-weights document rankings by multiplying base relevance scores with normalized user ratings (avg_relevance / 3). Additionally, fine_tune_index concatenates high-quality feedback entries with the original corpus and re-encodes them into a new FAISS index, effectively treating successful past responses as additional training data while keeping the embedding model frozen.
What format does the feedback data use and where is it stored?
Feedback is stored as newline-delimited JSON in data/feedback_data.json. Each entry contains numeric relevance and quality scores (1-5 scale), the original query, the generated response, and optional comments. The store_feedback function in retrieval_with_feedback_loop.py handles serialization, appending each feedback object as a new line to enable streaming reads and prevent data corruption during concurrent writes.
Can this implementation work with vector databases other than FAISS?
Yes, the architecture is database-agnostic at the interface level. While the reference implementation uses FAISS via encode_from_string in helper_functions.py, the RetrievalAugmentedGeneration class interacts with the vector store through LangChain's standard retriever interface (vectorstore.as_retriever()). You can substitute FAISS with Pinecone, Weaviate, or Chroma by replacing the initialization logic while preserving the adjust_relevance_scores and fine_tune_index methods that operate on the retrieved documents' metadata.
What threshold determines which feedback gets added to the index?
Feedback entries must meet a dual-threshold criteria of relevance ≥ 4 and quality ≥ 4 (on a 1-5 scale) to qualify for index fine-tuning. The fine_tune_index method filters the feedback dataset using these thresholds, then concatenates the qualifying entries with the original document corpus before re-encoding. This ensures only high-confidence, high-quality interactions augment the vector store, preventing noise from low-rated responses from degrading retrieval performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →