How to Process Documents for a Vector Database: A Complete Guide Using awesome-llm-apps

Processing documents for a vector database involves loading files with LangChain loaders, splitting text into overlapping chunks using RecursiveCharacterTextSplitter, and embedding them into a vector store like Qdrant using OpenAIEmbeddings.

The awesome-llm-apps repository by Shubhamsaboo demonstrates a production-ready pipeline for converting raw documents into searchable vectors. This guide breaks down the canonical pattern implemented across multiple RAG tutorials: load → split → embed → store, then retrieve → augment → answer.

The Three-Stage Document Processing Pipeline

The repository implements a consistent three-stage pipeline in rag_tutorials/rag_database_routing/rag_database_routing.py and related files. Each stage handles a specific transformation required to prepare documents for semantic search.

Stage 1: Loading Documents with PyPDFLoader

The first step converts binary file uploads into LangChain Document objects. The process_document helper function (lines 112-131 in rag_database_routing.py) handles this by writing uploaded bytes to a temporary file and invoking PyPDFLoader:

import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain_core.documents import Document
from typing import List

def process_document(file) -> List[Document]:
    # Write upload to a temp file

    with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
        tmp.write(file.getvalue())
        tmp_path = tmp.name

    # Load PDF into Document objects

    loader = PyPDFLoader(tmp_path)
    docs = loader.load()
    
    return docs

This pattern appears consistently across the repository, including in rag_tutorials/rag_agent_cohere/rag_agent_cohere.py and advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py.

Stage 2: Chunking with RecursiveCharacterTextSplitter

Large documents must be divided into smaller segments that fit within the embedding model's context window. The repository uses RecursiveCharacterTextSplitter configured with chunk_size=1000 and chunk_overlap=200 (lines 25-31 in rag_database_routing.py):

from langchain.text_splitter import RecursiveCharacterTextSplitter

def process_document(file) -> List[Document]:
    # ... loading logic ...

    docs = loader.load()
    
    # Split into overlapping chunks

    splitter = RecursiveCharacterTextSplitter(
        chunk_size=1000, 
        chunk_overlap=200
    )
    chunks = splitter.split_documents(docs)
    
    return chunks

The chunk_overlap parameter ensures semantic continuity between chunks, preventing context loss at segment boundaries. This configuration is replicated in rag_tutorials/hybrid_search_rag/main.py and the local hybrid search implementation.

Stage 3: Embedding and Storing in Qdrant

The final stage converts text chunks into vector embeddings and persists them to a vector store. The repository primarily uses Qdrant with OpenAI Embeddings (text-embedding-3-small, 1536 dimensions, cosine distance).

Collection initialization (lines 90-99 in rag_database_routing.py):

from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
from langchain_openai import OpenAIEmbeddings

client = QdrantClient(url="https://your-instance.com", api_key="YOUR_KEY")

# Create collection with OpenAI embedding dimensions

client.create_collection(
    collection_name="document_store",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)

vector_store = Qdrant(
    client=client,
    collection_name="document_store",
    embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)

Document insertion (lines 51-61 in rag_database_routing.py):


# Process multiple uploads

all_chunks = []
for file in uploaded_files:
    chunks = process_document(file)
    all_chunks.extend(chunks)

# Embed and store

vector_store.add_documents(all_chunks)

Complete Implementation: From Upload to Vector Store

Combining these stages into a production workflow requires handling file uploads, processing loops, and error handling. The repository implements this via Streamlit interfaces in rag_database_routing.py.

Full processing pipeline:

import streamlit as st
import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Qdrant
from langchain_openai import OpenAIEmbeddings
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

def process_document(file):
    """Load PDF and split into chunks."""
    with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
        tmp.write(file.getvalue())
        tmp_path = tmp.name
    
    loader = PyPDFLoader(tmp_path)
    docs = loader.load()
    
    splitter = RecursiveCharacterTextSplitter(
        chunk_size=1000, 
        chunk_overlap=200
    )
    return splitter.split_documents(docs)

# Initialize Qdrant

client = QdrantClient(url="https://your-instance.com", api_key="KEY")
client.create_collection(
    collection_name="docs",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)

vector_store = Qdrant(
    client=client,
    collection_name="docs",
    embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)

# Streamlit upload interface

uploaded = st.file_uploader("Upload PDFs", type="pdf", accept_multiple_files=True)
if uploaded:
    all_chunks = []
    for f in uploaded:
        all_chunks.extend(process_document(f))
    vector_store.add_documents(all_chunks)
    st.success(f"Indexed {len(all_chunks)} chunks from {len(uploaded)} files")

Retrieval and Querying

Once documents are processed into the vector database, the repository implements retrieval using LangChain's as_retriever interface. In rag_database_routing.py (lines 28-37), the system configures similarity search with a top-k of 4:

retriever = vector_store.as_retriever(
    search_type="similarity", 
    search_kwargs={"k": 4}
)

The retrieved chunks feed into a create_stuff_documents_chain combined with create_retrieval_chain to generate answers using the stored context. This pattern appears consistently across the RAG tutorials, including the Cohere implementation in rag_agent_cohere.py (lines 95-107).

Key Files and Variations

The document processing pipeline appears in multiple contexts throughout the repository:

File Role Key Variation
rag_tutorials/rag_database_routing/rag_database_routing.py Primary implementation with database routing logic Uses Qdrant with multiple collection selection based on similarity scores
rag_tutorials/rag_agent_cohere/rag_agent_cohere.py Cohere-specific RAG implementation Swaps OpenAI embeddings for Cohere embeddings while keeping the same chunking logic
rag_tutorials/local_hybrid_search_rag/local_main.py Local filesystem processing Processes documents from disk rather than Streamlit uploads
advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py Multi-agent legal application Uses the same process_document pattern for legal document analysis

Summary

Processing documents for a vector database in the awesome-llm-apps repository follows a modular, three-stage pipeline:

  • Load: Use PyPDFLoader (or similar LangChain loaders) to convert binary files into Document objects via temporary file handling
  • Split: Apply RecursiveCharacterTextSplitter with chunk_size=1000 and chunk_overlap=200 to create semantically coherent segments
  • Embed & Store: Initialize Qdrant with VectorParams(size=1536, distance=Distance.COSINE) for OpenAI embeddings, then persist chunks using vector_store.add_documents()

This architecture allows seamless swapping of components—whether changing embedding providers (Cohere vs. OpenAI), vector stores (Chroma vs. Qdrant), or document sources (PDF vs. text files)—while maintaining the core processing logic.

Frequently Asked Questions

What is the optimal chunk size for processing documents in a vector database?

The awesome-llm-apps repository uses chunk_size=1000 with chunk_overlap=200 as the default configuration in rag_database_routing.py. This balances context preservation with embedding model constraints, though you should adjust based on your specific embedding model's context window and the semantic structure of your documents.

Can I process file formats other than PDF using this pipeline?

Yes. While the repository primarily demonstrates PyPDFLoader for PDF processing, the process_document pattern can accommodate any LangChain document loader. For text files, substitute TextLoader; for CSV data, use CSVLoader; for web content, use WebBaseLoader. The chunking and embedding stages remain identical regardless of the source format.

How do I choose between Qdrant and other vector stores like Chroma or FAISS?

The repository implements Qdrant in rag_database_routing.py due to its production-ready features and LangChain integration. However, the vector store is abstracted via LangChain's VectorStore interface, allowing you to substitute Chroma for local development, FAISS for in-memory experimentation, or Pinecone for managed cloud infrastructure without changing the document processing logic.

What embedding model should I use when processing documents for production?

According to the source code in rag_database_routing.py, the repository defaults to OpenAIEmbeddings(model="text-embedding-3-small") with 1536 dimensions and cosine distance. For production, text-embedding-3-small offers an optimal balance of cost and performance, though the Cohere tutorial in rag_agent_cohere.py demonstrates how to swap in CohereEmbeddings if you require a different provider or multilingual support.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →