How to Implement Basic RAG with LangChain: A Step-by-Step Guide
You can implement basic RAG with LangChain by loading documents with PyPDFLoader, splitting text using RecursiveCharacterTextSplitter, storing embeddings in a FAISS vector store, and retrieving relevant chunks with as_retriever() to augment your LLM queries.
Implementing basic RAG with LangChain provides the foundational architecture for retrieval-augmented generation without complex orchestration frameworks. The NirDiamant/RAG_Techniques repository demonstrates this pattern using pure Python components that chain together document processing, embedding, and retrieval stages.
The Four Stages of a Basic RAG Pipeline
A minimal RAG workflow built with LangChain consists of four logical stages, each implemented with specific components in the repository.
Stage 1: Document Loading with PyPDFLoader
The pipeline begins by ingesting raw documents. In helper_functions.py, the PyPDFLoader class from LangChain community loaders handles PDF parsing.
from langchain_community.document_loaders import PyPDFLoader
# Located at line 1 in helper_functions.py
loader = PyPDFLoader(path)
documents = loader.load()
This stage extracts text from each page of the PDF, creating a list of Document objects that preserve page metadata.
Stage 2: Text Chunking with RecursiveCharacterTextSplitter
Long documents must be divided into manageable pieces to fit within embedding model context windows. The repository uses RecursiveCharacterTextSplitter at lines 65-71 in helper_functions.py.
from langchain.text_splitter import RecursiveCharacterTextSplitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size, # Default: 1000
chunk_overlap=chunk_overlap, # Default: 200
separators=["\n\n", "\n", ".", "!", "?", ",", " ", ""]
)
chunks = text_splitter.split_documents(documents)
This splitter recursively divides text at natural boundaries, ensuring semantically coherent chunks while respecting the specified character limits.
Stage 3: Embedding and Vector Storage with FAISS
Each chunk is converted to a dense vector embedding and indexed for fast similarity search. The encode_pdf function (lines 48-76 in helper_functions.py) orchestrates this using OpenAIEmbeddings and FAISS.
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
def encode_pdf(path, chunk_size=1000, chunk_overlap=200):
# Load and split
loader = PyPDFLoader(path)
documents = loader.load()
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=chunk_size, chunk_overlap=chunk_overlap
)
texts = text_splitter.split_documents(documents)
# Clean tabs for better embedding quality
for text in texts:
text.page_content = text.page_content.replace('\t', ' ')
# Create embeddings and vector store
embeddings = OpenAIEmbeddings()
vector_store = FAISS.from_documents(texts, embeddings)
return vector_store
FAISS provides in-memory vector storage with efficient similarity search, making it ideal for prototyping without external database dependencies.
Stage 4: Retrieval and Context Generation
The final stage retrieves relevant chunks for a user query. In simple_rag.py (lines 41-57), the pipeline creates a retriever and fetches context.
from helper_functions import retrieve_context_per_question
# Create retriever with top-k configuration
retriever = vector_store.as_retriever(search_kwargs={"k": n_retrieved})
# Retrieve relevant chunks
chunks = retrieve_context_per_question(query, retriever)
The as_retriever() method converts the FAISS store into a LangChain retriever interface, allowing seamless integration with chains and agents.
Complete Implementation Example
Below is a self-contained function that reproduces the basic RAG pipeline from the repository. This mirrors the logic in simple_rag.py but packages it as a reusable utility.
import os
from dotenv import load_dotenv
from helper_functions import encode_pdf, retrieve_context_per_question
# Load environment variables
load_dotenv()
os.environ["OPENAI_API_KEY"] = os.getenv("OPENAI_API_KEY")
def basic_rag(
pdf_path: str,
query: str,
chunk_size: int = 1000,
chunk_overlap: int = 200,
n_retrieved: int = 2,
):
"""
End-to-end basic RAG implementation:
1. Encode PDF into FAISS vector store
2. Retrieve top-k relevant chunks for the query
Returns:
List of retrieved text chunks
"""
# Build vector store from PDF
vector_store = encode_pdf(
path=pdf_path,
chunk_size=chunk_size,
chunk_overlap=chunk_overlap,
)
# Configure retriever
retriever = vector_store.as_retriever(
search_kwargs={"k": n_retrieved}
)
# Retrieve relevant context
chunks = retrieve_context_per_question(query, retriever)
return chunks
# Example usage
if __name__ == "__main__":
pdf_file = "../data/Understanding_Climate_Change.pdf"
user_question = "What is the main cause of climate change?"
result_chunks = basic_rag(pdf_file, user_question)
print("\n--- Retrieved Chunks ---")
for i, chunk in enumerate(result_chunks, 1):
print(f"\nChunk {i}:\n{chunk[:500]}...")
Extending to LLM-Generated Answers
To generate natural language answers rather than returning raw chunks, extend the pipeline with a LangChain LLM chain. The repository provides create_question_answer_from_context_chain in helper_functions.py for this purpose.
from helper_functions import create_question_answer_from_context_chain, answer_question_from_context
from langchain_openai import ChatOpenAI
# Initialize LLM with deterministic output
llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0)
# Build the Q&A chain using the repository's prompt template
qa_chain = create_question_answer_from_context_chain(llm)
# Concatenate retrieved chunks
context = "\n".join(chunks)
# Generate answer
answer_dict = answer_question_from_context(user_question, context, qa_chain)
print(f"Answer: {answer_dict['answer']}")
This composition follows the standard LangChain pattern of retrieval-augmented generation, where the retrieved context is injected into a prompt template to ground the LLM's response in source documents.
Key Files in the Repository
The NirDiamant/RAG_Techniques repository organizes the basic RAG implementation across the following files:
-
helper_functions.py– Core utilities includingencode_pdf()(lines 48-76) for vector store creation,retrieve_context_per_question()for retrieval logic, andcreate_question_answer_from_context_chain()for LLM answer generation. -
all_rag_techniques_runnable_scripts/simple_rag.py– End-to-end runnable script demonstrating the complete pipeline from environment setup through retrieval (lines 41-57 contain the core retrieval logic). -
evaluation/evalute_rag.py– Evaluation utilities for measuring retrieval precision and recall, useful for benchmarking your basic RAG setup. -
data/Understanding_Climate_Change.pdf– Sample document used in examples; any PDF can be substituted.
Summary
-
Implement basic RAG with LangChain by chaining four components:
PyPDFLoaderfor ingestion,RecursiveCharacterTextSplitterfor chunking,OpenAIEmbeddingswithFAISSfor vector storage, andas_retriever()for similarity search. -
The
encode_pdf()function inhelper_functions.pyencapsulates the indexing pipeline, whilesimple_rag.pydemonstrates the retrieval execution. -
FAISS provides in-memory vector storage suitable for prototyping, requiring only the OpenAI API key for embeddings.
-
Extend the basic retrieval pipeline to full generation by composing
ChatOpenAIwithcreate_question_answer_from_context_chain()to produce grounded answers from retrieved chunks.
Frequently Asked Questions
What is the minimum code needed to implement basic RAG with LangChain?
The minimal implementation requires four lines of core logic: load documents with PyPDFLoader, split with RecursiveCharacterTextSplitter, create a FAISS vector store with OpenAIEmbeddings, and call as_retriever() to fetch relevant chunks. The basic_rag() function in the examples above encapsulates this into a reusable 15-line utility.
How does the FAISS vector store work in this implementation?
FAISS (Facebook AI Similarity Search) serves as the in-memory vector database that stores document embeddings created by OpenAIEmbeddings. When you call vector_store.as_retriever(search_kwargs={"k": n}), LangChain queries the FAISS index using cosine similarity to return the top-n most relevant chunks without requiring an external database service.
Can I use a different embedding model instead of OpenAIEmbeddings?
Yes, you can swap OpenAIEmbeddings for alternatives like HuggingFaceEmbeddings, CohereEmbeddings, or BedrockEmbeddings by modifying the encode_pdf() function in helper_functions.py. The repository includes get_langchain_embedding_provider() to facilitate swapping providers while maintaining the same interface for vector store creation.
How do I evaluate the quality of my RAG retrieval?
Use the evaluation/evalute_rag.py module in the repository, which provides metrics for precision, recall, and mean reciprocal rank (MRR). You can benchmark your retriever by comparing the retrieved chunks against a ground-truth dataset of relevant passages for specific queries, allowing you to tune parameters like chunk_size and n_retrieved for optimal performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →