How to Process Documents for a Vector Database: A Complete Guide Using awesome-llm-apps
Processing documents for a vector database involves loading files with LangChain loaders, splitting text into overlapping chunks using RecursiveCharacterTextSplitter, and embedding them into a vector store like Qdrant using OpenAIEmbeddings.
The awesome-llm-apps repository by Shubhamsaboo demonstrates a production-ready pipeline for converting raw documents into searchable vectors. This guide breaks down the canonical pattern implemented across multiple RAG tutorials: load → split → embed → store, then retrieve → augment → answer.
The Three-Stage Document Processing Pipeline
The repository implements a consistent three-stage pipeline in rag_tutorials/rag_database_routing/rag_database_routing.py and related files. Each stage handles a specific transformation required to prepare documents for semantic search.
Stage 1: Loading Documents with PyPDFLoader
The first step converts binary file uploads into LangChain Document objects. The process_document helper function (lines 112-131 in rag_database_routing.py) handles this by writing uploaded bytes to a temporary file and invoking PyPDFLoader:
import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain_core.documents import Document
from typing import List
def process_document(file) -> List[Document]:
# Write upload to a temp file
with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
tmp.write(file.getvalue())
tmp_path = tmp.name
# Load PDF into Document objects
loader = PyPDFLoader(tmp_path)
docs = loader.load()
return docs
This pattern appears consistently across the repository, including in rag_tutorials/rag_agent_cohere/rag_agent_cohere.py and advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py.
Stage 2: Chunking with RecursiveCharacterTextSplitter
Large documents must be divided into smaller segments that fit within the embedding model's context window. The repository uses RecursiveCharacterTextSplitter configured with chunk_size=1000 and chunk_overlap=200 (lines 25-31 in rag_database_routing.py):
from langchain.text_splitter import RecursiveCharacterTextSplitter
def process_document(file) -> List[Document]:
# ... loading logic ...
docs = loader.load()
# Split into overlapping chunks
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200
)
chunks = splitter.split_documents(docs)
return chunks
The chunk_overlap parameter ensures semantic continuity between chunks, preventing context loss at segment boundaries. This configuration is replicated in rag_tutorials/hybrid_search_rag/main.py and the local hybrid search implementation.
Stage 3: Embedding and Storing in Qdrant
The final stage converts text chunks into vector embeddings and persists them to a vector store. The repository primarily uses Qdrant with OpenAI Embeddings (text-embedding-3-small, 1536 dimensions, cosine distance).
Collection initialization (lines 90-99 in rag_database_routing.py):
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
from langchain_openai import OpenAIEmbeddings
client = QdrantClient(url="https://your-instance.com", api_key="YOUR_KEY")
# Create collection with OpenAI embedding dimensions
client.create_collection(
collection_name="document_store",
vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)
vector_store = Qdrant(
client=client,
collection_name="document_store",
embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)
Document insertion (lines 51-61 in rag_database_routing.py):
# Process multiple uploads
all_chunks = []
for file in uploaded_files:
chunks = process_document(file)
all_chunks.extend(chunks)
# Embed and store
vector_store.add_documents(all_chunks)
Complete Implementation: From Upload to Vector Store
Combining these stages into a production workflow requires handling file uploads, processing loops, and error handling. The repository implements this via Streamlit interfaces in rag_database_routing.py.
Full processing pipeline:
import streamlit as st
import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Qdrant
from langchain_openai import OpenAIEmbeddings
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
def process_document(file):
"""Load PDF and split into chunks."""
with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
tmp.write(file.getvalue())
tmp_path = tmp.name
loader = PyPDFLoader(tmp_path)
docs = loader.load()
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200
)
return splitter.split_documents(docs)
# Initialize Qdrant
client = QdrantClient(url="https://your-instance.com", api_key="KEY")
client.create_collection(
collection_name="docs",
vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)
vector_store = Qdrant(
client=client,
collection_name="docs",
embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)
# Streamlit upload interface
uploaded = st.file_uploader("Upload PDFs", type="pdf", accept_multiple_files=True)
if uploaded:
all_chunks = []
for f in uploaded:
all_chunks.extend(process_document(f))
vector_store.add_documents(all_chunks)
st.success(f"Indexed {len(all_chunks)} chunks from {len(uploaded)} files")
Retrieval and Querying
Once documents are processed into the vector database, the repository implements retrieval using LangChain's as_retriever interface. In rag_database_routing.py (lines 28-37), the system configures similarity search with a top-k of 4:
retriever = vector_store.as_retriever(
search_type="similarity",
search_kwargs={"k": 4}
)
The retrieved chunks feed into a create_stuff_documents_chain combined with create_retrieval_chain to generate answers using the stored context. This pattern appears consistently across the RAG tutorials, including the Cohere implementation in rag_agent_cohere.py (lines 95-107).
Key Files and Variations
The document processing pipeline appears in multiple contexts throughout the repository:
| File | Role | Key Variation |
|---|---|---|
rag_tutorials/rag_database_routing/rag_database_routing.py |
Primary implementation with database routing logic | Uses Qdrant with multiple collection selection based on similarity scores |
rag_tutorials/rag_agent_cohere/rag_agent_cohere.py |
Cohere-specific RAG implementation | Swaps OpenAI embeddings for Cohere embeddings while keeping the same chunking logic |
rag_tutorials/local_hybrid_search_rag/local_main.py |
Local filesystem processing | Processes documents from disk rather than Streamlit uploads |
advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py |
Multi-agent legal application | Uses the same process_document pattern for legal document analysis |
Summary
Processing documents for a vector database in the awesome-llm-apps repository follows a modular, three-stage pipeline:
- Load: Use
PyPDFLoader(or similar LangChain loaders) to convert binary files intoDocumentobjects via temporary file handling - Split: Apply
RecursiveCharacterTextSplitterwithchunk_size=1000andchunk_overlap=200to create semantically coherent segments - Embed & Store: Initialize Qdrant with
VectorParams(size=1536, distance=Distance.COSINE)for OpenAI embeddings, then persist chunks usingvector_store.add_documents()
This architecture allows seamless swapping of components—whether changing embedding providers (Cohere vs. OpenAI), vector stores (Chroma vs. Qdrant), or document sources (PDF vs. text files)—while maintaining the core processing logic.
Frequently Asked Questions
What is the optimal chunk size for processing documents in a vector database?
The awesome-llm-apps repository uses chunk_size=1000 with chunk_overlap=200 as the default configuration in rag_database_routing.py. This balances context preservation with embedding model constraints, though you should adjust based on your specific embedding model's context window and the semantic structure of your documents.
Can I process file formats other than PDF using this pipeline?
Yes. While the repository primarily demonstrates PyPDFLoader for PDF processing, the process_document pattern can accommodate any LangChain document loader. For text files, substitute TextLoader; for CSV data, use CSVLoader; for web content, use WebBaseLoader. The chunking and embedding stages remain identical regardless of the source format.
How do I choose between Qdrant and other vector stores like Chroma or FAISS?
The repository implements Qdrant in rag_database_routing.py due to its production-ready features and LangChain integration. However, the vector store is abstracted via LangChain's VectorStore interface, allowing you to substitute Chroma for local development, FAISS for in-memory experimentation, or Pinecone for managed cloud infrastructure without changing the document processing logic.
What embedding model should I use when processing documents for production?
According to the source code in rag_database_routing.py, the repository defaults to OpenAIEmbeddings(model="text-embedding-3-small") with 1536 dimensions and cosine distance. For production, text-embedding-3-small offers an optimal balance of cost and performance, though the Cohere tutorial in rag_agent_cohere.py demonstrates how to swap in CohereEmbeddings if you require a different provider or multilingual support.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →