# How to Process Documents for a Vector Database: A Complete Guide Using awesome-llm-apps

> Learn to process documents for a vector database. This guide shows how to load files, split text, and embed into Qdrant using LangChain and OpenAIEmbeddings from awesome-llm-apps.

- Repository: [Shubham Saboo/awesome-llm-apps](https://github.com/shubhamsaboo/awesome-llm-apps)
- Tags: how-to-guide
- Published: 2026-02-19

---

**Processing documents for a vector database involves loading files with LangChain loaders, splitting text into overlapping chunks using RecursiveCharacterTextSplitter, and embedding them into a vector store like Qdrant using OpenAIEmbeddings.**

The `awesome-llm-apps` repository by Shubhamsaboo demonstrates a production-ready pipeline for converting raw documents into searchable vectors. This guide breaks down the canonical pattern implemented across multiple RAG tutorials: **load → split → embed → store**, then **retrieve → augment → answer**.

## The Three-Stage Document Processing Pipeline

The repository implements a consistent three-stage pipeline in [`rag_tutorials/rag_database_routing/rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_database_routing/rag_database_routing.py) and related files. Each stage handles a specific transformation required to prepare documents for semantic search.

### Stage 1: Loading Documents with PyPDFLoader

The first step converts binary file uploads into LangChain `Document` objects. The `process_document` helper function (lines 112-131 in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py)) handles this by writing uploaded bytes to a temporary file and invoking `PyPDFLoader`:

```python
import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain_core.documents import Document
from typing import List

def process_document(file) -> List[Document]:
    # Write upload to a temp file

    with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
        tmp.write(file.getvalue())
        tmp_path = tmp.name

    # Load PDF into Document objects

    loader = PyPDFLoader(tmp_path)
    docs = loader.load()
    
    return docs

```

This pattern appears consistently across the repository, including in [`rag_tutorials/rag_agent_cohere/rag_agent_cohere.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_agent_cohere/rag_agent_cohere.py) and [`advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py).

### Stage 2: Chunking with RecursiveCharacterTextSplitter

Large documents must be divided into smaller segments that fit within the embedding model's context window. The repository uses `RecursiveCharacterTextSplitter` configured with `chunk_size=1000` and `chunk_overlap=200` (lines 25-31 in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py)):

```python
from langchain.text_splitter import RecursiveCharacterTextSplitter

def process_document(file) -> List[Document]:
    # ... loading logic ...

    docs = loader.load()
    
    # Split into overlapping chunks

    splitter = RecursiveCharacterTextSplitter(
        chunk_size=1000, 
        chunk_overlap=200
    )
    chunks = splitter.split_documents(docs)
    
    return chunks

```

The `chunk_overlap` parameter ensures semantic continuity between chunks, preventing context loss at segment boundaries. This configuration is replicated in [`rag_tutorials/hybrid_search_rag/main.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/hybrid_search_rag/main.py) and the local hybrid search implementation.

### Stage 3: Embedding and Storing in Qdrant

The final stage converts text chunks into vector embeddings and persists them to a vector store. The repository primarily uses **Qdrant** with **OpenAI Embeddings** (`text-embedding-3-small`, 1536 dimensions, cosine distance).

Collection initialization (lines 90-99 in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py)):

```python
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
from langchain_openai import OpenAIEmbeddings

client = QdrantClient(url="https://your-instance.com", api_key="YOUR_KEY")

# Create collection with OpenAI embedding dimensions

client.create_collection(
    collection_name="document_store",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)

vector_store = Qdrant(
    client=client,
    collection_name="document_store",
    embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)

```

Document insertion (lines 51-61 in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py)):

```python

# Process multiple uploads

all_chunks = []
for file in uploaded_files:
    chunks = process_document(file)
    all_chunks.extend(chunks)

# Embed and store

vector_store.add_documents(all_chunks)

```

## Complete Implementation: From Upload to Vector Store

Combining these stages into a production workflow requires handling file uploads, processing loops, and error handling. The repository implements this via Streamlit interfaces in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py).

**Full processing pipeline:**

```python
import streamlit as st
import tempfile
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Qdrant
from langchain_openai import OpenAIEmbeddings
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

def process_document(file):
    """Load PDF and split into chunks."""
    with tempfile.NamedTemporaryFile(delete=False, suffix=".pdf") as tmp:
        tmp.write(file.getvalue())
        tmp_path = tmp.name
    
    loader = PyPDFLoader(tmp_path)
    docs = loader.load()
    
    splitter = RecursiveCharacterTextSplitter(
        chunk_size=1000, 
        chunk_overlap=200
    )
    return splitter.split_documents(docs)

# Initialize Qdrant

client = QdrantClient(url="https://your-instance.com", api_key="KEY")
client.create_collection(
    collection_name="docs",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE)
)

vector_store = Qdrant(
    client=client,
    collection_name="docs",
    embeddings=OpenAIEmbeddings(model="text-embedding-3-small")
)

# Streamlit upload interface

uploaded = st.file_uploader("Upload PDFs", type="pdf", accept_multiple_files=True)
if uploaded:
    all_chunks = []
    for f in uploaded:
        all_chunks.extend(process_document(f))
    vector_store.add_documents(all_chunks)
    st.success(f"Indexed {len(all_chunks)} chunks from {len(uploaded)} files")

```

## Retrieval and Querying

Once documents are processed into the vector database, the repository implements retrieval using LangChain's `as_retriever` interface. In [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py) (lines 28-37), the system configures similarity search with a top-k of 4:

```python
retriever = vector_store.as_retriever(
    search_type="similarity", 
    search_kwargs={"k": 4}
)

```

The retrieved chunks feed into a `create_stuff_documents_chain` combined with `create_retrieval_chain` to generate answers using the stored context. This pattern appears consistently across the RAG tutorials, including the Cohere implementation in [`rag_agent_cohere.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_agent_cohere.py) (lines 95-107).

## Key Files and Variations

The document processing pipeline appears in multiple contexts throughout the repository:

| File | Role | Key Variation |
|------|------|---------------|
| [`rag_tutorials/rag_database_routing/rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_database_routing/rag_database_routing.py) | Primary implementation with database routing logic | Uses Qdrant with multiple collection selection based on similarity scores |
| [`rag_tutorials/rag_agent_cohere/rag_agent_cohere.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/rag_agent_cohere/rag_agent_cohere.py) | Cohere-specific RAG implementation | Swaps OpenAI embeddings for Cohere embeddings while keeping the same chunking logic |
| [`rag_tutorials/local_hybrid_search_rag/local_main.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/local_hybrid_search_rag/local_main.py) | Local filesystem processing | Processes documents from disk rather than Streamlit uploads |
| [`advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/advanced_ai_agents/multi_agent_apps/agent_teams/ai_legal_agent_team/local_ai_legal_agent_team/local_legal_agent.py) | Multi-agent legal application | Uses the same `process_document` pattern for legal document analysis |

## Summary

Processing documents for a vector database in the `awesome-llm-apps` repository follows a modular, three-stage pipeline:

- **Load**: Use `PyPDFLoader` (or similar LangChain loaders) to convert binary files into `Document` objects via temporary file handling
- **Split**: Apply `RecursiveCharacterTextSplitter` with `chunk_size=1000` and `chunk_overlap=200` to create semantically coherent segments
- **Embed & Store**: Initialize Qdrant with `VectorParams(size=1536, distance=Distance.COSINE)` for OpenAI embeddings, then persist chunks using `vector_store.add_documents()`

This architecture allows seamless swapping of components—whether changing embedding providers (Cohere vs. OpenAI), vector stores (Chroma vs. Qdrant), or document sources (PDF vs. text files)—while maintaining the core processing logic.

## Frequently Asked Questions

### What is the optimal chunk size for processing documents in a vector database?

The `awesome-llm-apps` repository uses `chunk_size=1000` with `chunk_overlap=200` as the default configuration in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py). This balances context preservation with embedding model constraints, though you should adjust based on your specific embedding model's context window and the semantic structure of your documents.

### Can I process file formats other than PDF using this pipeline?

Yes. While the repository primarily demonstrates `PyPDFLoader` for PDF processing, the `process_document` pattern can accommodate any LangChain document loader. For text files, substitute `TextLoader`; for CSV data, use `CSVLoader`; for web content, use `WebBaseLoader`. The chunking and embedding stages remain identical regardless of the source format.

### How do I choose between Qdrant and other vector stores like Chroma or FAISS?

The repository implements Qdrant in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py) due to its production-ready features and LangChain integration. However, the vector store is abstracted via LangChain's `VectorStore` interface, allowing you to substitute Chroma for local development, FAISS for in-memory experimentation, or Pinecone for managed cloud infrastructure without changing the document processing logic.

### What embedding model should I use when processing documents for production?

According to the source code in [`rag_database_routing.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_database_routing.py), the repository defaults to `OpenAIEmbeddings(model="text-embedding-3-small")` with 1536 dimensions and cosine distance. For production, `text-embedding-3-small` offers an optimal balance of cost and performance, though the Cohere tutorial in [`rag_agent_cohere.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_agent_cohere.py) demonstrates how to swap in `CohereEmbeddings` if you require a different provider or multilingual support.