Implementing RAG with Local Open-Source LLMs: A Complete Privacy-First Guide
You can build fully private Retrieval-Augmented Generation (RAG) pipelines using local open-source LLMs like Llama 3.1 and Gemma by running the entire stack—embeddings, vector storage, and inference—on your own machine via Ollama, eliminating all external API calls and data exposure.
This guide walks through production-ready implementations from the awesome-llm-apps repository that demonstrate how to process sensitive documents without sending data to third-party servers. By leveraging frameworks like LangChain and Agno alongside local vector stores such as Chroma and LanceDB, you can create privacy-preserving AI assistants that keep your proprietary information strictly on-device.
Why Local RAG Matters for Privacy
Traditional RAG implementations rely on cloud-based embedding models and LLM APIs, which require transmitting your documents and queries to external servers. For organizations handling confidential legal contracts, medical records, or proprietary research, this creates unacceptable data sovereignty risks.
Implementing RAG with local open-source LLMs solves this by ensuring:
- Zero network egress: All text processing occurs within your local network
- Model ownership: You control the weights and can audit the inference code
- Compliance: Easier alignment with GDPR, HIPAA, and SOC 2 requirements that restrict data residency
Architecture Overview
The awesome-llm-apps repository provides two distinct architectural patterns for local RAG, both following the standard ingest-chunk-embed-retrieve-generate flow while keeping every step local.
Data Ingestion and Chunking
Documents enter the pipeline through loaders like WebBaseLoader (for URLs) or direct PDF parsers. The RecursiveCharacterTextSplitter breaks documents into 500-character chunks with 10-character overlaps, preserving semantic boundaries while creating manageable pieces for the context window.
Local Embedding Generation
Instead of calling OpenAI or Cohere APIs, both implementations use Ollama-hosted embedding models:
- Llama 3.1 pipeline: Uses
OllamaEmbeddingsfrom LangChain withmodel="llama3.1" - Gemma pipeline: Uses
OllamaEmbedderfrom Agno withid="embeddinggemma:latest"
These models run on your local Ollama server at http://127.0.0.1:11434, converting text chunks into vector representations without network transmission.
Vector Storage Options
The repository demonstrates two lightweight local vector databases:
- Chroma: In-memory storage ideal for session-based RAG with Llama 3.1
- LanceDB: File-based persistent storage used with Gemma embeddings, allowing knowledge bases to survive across application restarts
Local LLM Inference
Generation stays local using Ollama-hosted chat models:
ChatOllama(LangChain) for the Llama 3.1 implementationOllama(Agno) for the Gemma-based agentic RAG
Both interfaces stream tokens directly from your local GPU or CPU, ensuring prompts and responses never leave the machine.
Implementation 1: Llama 3.1 with LangChain and Chroma
The first implementation in rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py demonstrates a straightforward RAG pipeline for querying web pages and documents using Meta's Llama 3.1 model.
# File: rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py
# https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py
import streamlit as st
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import WebBaseLoader
from langchain_community.vectorstores import Chroma
from langchain_ollama import OllamaEmbeddings, ChatOllama
# 1. UI – ask for a webpage URL
webpage_url = st.text_input("Enter Webpage URL")
if webpage_url:
# 2. Load & split the page
docs = WebBaseLoader(webpage_url).load()
splits = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=10
).split_documents(docs)
# 3. Create local embeddings + vector store
embeddings = OllamaEmbeddings(
model="llama3.1",
base_url="http://127.0.0.1:11434"
)
vectorstore = Chroma.from_documents(splits, embeddings)
# 4. LLM (local Ollama)
ollama = ChatOllama(
model="llama3.1",
base_url="http://127.0.0.1:11434"
)
def answer(question):
# Retrieve most-relevant chunks
docs = vectorstore.as_retriever().invoke(question)
context = "\n\n".join(d.page_content for d in docs)
# Prompt the model (local)
prompt = f"Question: {question}\n\nContext: {context}"
response = ollama.invoke([('human', prompt)])
return response.content.strip()
# 5. Ask a question about the page
query = st.text_input("Ask any question about the webpage")
if query:
st.write(answer(query))
Key implementation details:
- Local-only processing: The
base_url="http://127.0.0.1:11434"parameter ensures all embedding and generation requests route to your local Ollama instance. - In-memory vector store:
Chroma.from_documents()creates a temporary vector database that exists only for the session, perfect for ephemeral document analysis. - Simple retrieval: The
as_retriever().invoke()method performs semantic search against the local Chroma store, returning the top-k relevant chunks.
Implementation 2: Gemma Embeddings with Agno and LanceDB
The second implementation in rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py showcases an agentic approach using Google's Gemma for embeddings and LanceDB for persistent storage.
# File: rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py
# https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py
import streamlit as st
from agno.agent import Agent
from agno.knowledge.embedder.ollama import OllamaEmbedder
from agno.knowledge.knowledge import Knowledge
from agno.models.ollama import Ollama
from agno.vectordb.lancedb import LanceDb, SearchType
# 1. Set up a persistent knowledge base that uses the Gemma embedder
@st.cache_resource
def load_kb():
return Knowledge(
vector_db=LanceDb(
table_name="docs",
uri="tmp/lancedb",
search_type=SearchType.vector,
embedder=OllamaEmbedder(
id="embeddinggemma:latest",
dimensions=768
),
)
)
kb = load_kb()
# 2. UI – let the user add PDFs/URLs
new_url = st.sidebar.text_input(
"Add URL",
placeholder="https://example.com/file.pdf"
)
if st.sidebar.button("Add"):
if new_url and new_url not in st.session_state.get("urls", []):
st.session_state.setdefault("urls", []).append(new_url)
kb.add_content(url=new_url) # Gemma creates the embeddings locally
# 3. Build an Agno Agent that uses a local Llama 3.2 model for generation
agent = Agent(
model=Ollama(id="llama3.2:latest"),
knowledge=kb,
instructions=[
"Search the knowledge base for relevant info and answer based on it.",
"Use clear headings, bullet points, and concise language."
],
search_knowledge=True,
markdown=True,
)
# 4. Ask a question – stream the answer back to the UI
question = st.text_input("Enter your question:")
if st.button("Get Answer") and question:
with st.spinner("Generating answer…"):
answer = ""
placeholder = st.empty()
for chunk in agent.run(question, stream=True):
answer += chunk.content or ""
placeholder.markdown(answer)
Key implementation details:
- Persistent vector storage:
LanceDbwithuri="tmp/lancedb"writes embeddings to disk, enabling the knowledge base to persist across application restarts without re-indexing. - Gemma embeddings: The
OllamaEmbedderwithid="embeddinggemma:latest"generates 768-dimensional vectors locally, providing high-quality semantic representations without cloud dependencies. - Agentic orchestration: The Agno
Agentclass handles retrieval and generation automatically whensearch_knowledge=True, streaming responses token-by-token for real-time UI updates.
Running the Applications Locally
To deploy these privacy-preserving RAG systems on your own hardware, follow these setup steps:
-
Install Ollama from ollama.com and verify the service is running on
http://127.0.0.1:11434. -
Pull the required models:
ollama pull llama3.1 ollama pull embeddinggemma:latest ollama pull llama3.2:latest -
Clone the repository and install dependencies:
git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git cd awesome-llm-apps pip install -r requirements.txt -
Launch the Streamlit applications:
# For Llama 3.1 RAG streamlit run rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py # For Gemma embedding RAG streamlit run rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py -
Interact via browser – input URLs or PDFs, then query the knowledge base. All embeddings, retrievals, and generations execute locally on your CPU or GPU.
Summary
Implementing RAG with local open-source LLMs enables organizations to leverage AI-powered document analysis while maintaining complete data sovereignty. The key takeaways from the awesome-llm-apps implementations include:
- Zero external dependencies: Both the Llama 3.1 and Gemma pipelines run entirely through Ollama on
http://127.0.0.1:11434, ensuring no data leaves your local network. - Flexible vector storage: Choose Chroma for ephemeral, in-memory sessions or LanceDB for persistent, file-based knowledge bases that survive application restarts.
- Modular architecture: Swapping models requires only changing the
modelparameter inOllamaEmbeddingsorOllamaEmbedder, allowing easy experimentation with new open-source releases. - Agentic capabilities: The Agno framework demonstrates how local RAG can evolve into autonomous agents that reason over private knowledge bases without external API costs.
Frequently Asked Questions
Can I use these local RAG implementations for commercial documents without violating data privacy regulations?
Yes. Because all processing—including embedding generation, vector storage, and LLM inference—occurs locally via Ollama, your documents never transmit to third-party servers. This architecture supports compliance with GDPR, HIPAA, and SOC 2 requirements that mandate data residency, though you should still verify that your local hardware and Ollama installation meet your organization's specific security standards.
How do I switch from Llama 3.1 to a different local model like Mistral or Qwen?
Changing the underlying LLM requires only updating the model identifier in your code. For the LangChain implementation in rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py, change the model parameter in both OllamaEmbeddings and ChatOllama to your preferred Ollama model (e.g., "mistral" or "qwen2.5"). For the Agno implementation, update the Ollama model ID and ensure you have pulled the corresponding model via ollama pull <model_name>.
What are the hardware requirements for running these privacy-preserving RAG systems?
The implementations run on consumer hardware, though performance varies by model size. Llama 3.1 (8B parameters) and Gemma embedding models operate comfortably on machines with 16GB RAM and a modern CPU, though a GPU with 8GB+ VRAM significantly improves inference speed. LanceDB requires minimal disk space (megabytes to gigabytes depending on corpus size), while Chroma operates entirely in memory, requiring sufficient RAM to hold your document chunks and vectors simultaneously.
Can I extend these implementations to handle images or multi-modal documents?
Yes. The Gemma embedding model (embeddinggemma:latest) supports multi-modal inputs, and the Agno framework's architecture in rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py can be extended to process images. You would replace the text-only loaders with libraries like PyMuPDF for PDF image extraction or Pillow for direct image loading, then pass the processed image data to the embedding model. The LanceDB vector store can store multi-modal embeddings alongside text vectors, enabling RAG over mixed document types while maintaining the same local privacy guarantees.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →