# Implementing RAG with Local Open-Source LLMs: A Complete Privacy-First Guide

> Build private RAG pipelines with local LLMs like Llama 3.1 and Gemma. Run embeddings, vector storage, and inference on your machine with Ollama for complete data privacy. No external APIs.

- Repository: [Shubham Saboo/awesome-llm-apps](https://github.com/shubhamsaboo/awesome-llm-apps)
- Tags: how-to-guide
- Published: 2026-02-16

---

**You can build fully private Retrieval-Augmented Generation (RAG) pipelines using local open-source LLMs like Llama 3.1 and Gemma by running the entire stack—embeddings, vector storage, and inference—on your own machine via Ollama, eliminating all external API calls and data exposure.**

This guide walks through production-ready implementations from the **awesome-llm-apps** repository that demonstrate how to process sensitive documents without sending data to third-party servers. By leveraging frameworks like LangChain and Agno alongside local vector stores such as Chroma and LanceDB, you can create privacy-preserving AI assistants that keep your proprietary information strictly on-device.

## Why Local RAG Matters for Privacy

Traditional RAG implementations rely on cloud-based embedding models and LLM APIs, which require transmitting your documents and queries to external servers. For organizations handling confidential legal contracts, medical records, or proprietary research, this creates unacceptable data sovereignty risks.

Implementing RAG with local open-source LLMs solves this by ensuring:
- **Zero network egress**: All text processing occurs within your local network
- **Model ownership**: You control the weights and can audit the inference code
- **Compliance**: Easier alignment with GDPR, HIPAA, and SOC 2 requirements that restrict data residency

## Architecture Overview

The awesome-llm-apps repository provides two distinct architectural patterns for local RAG, both following the standard ingest-chunk-embed-retrieve-generate flow while keeping every step local.

### Data Ingestion and Chunking

Documents enter the pipeline through loaders like `WebBaseLoader` (for URLs) or direct PDF parsers. The `RecursiveCharacterTextSplitter` breaks documents into 500-character chunks with 10-character overlaps, preserving semantic boundaries while creating manageable pieces for the context window.

### Local Embedding Generation

Instead of calling OpenAI or Cohere APIs, both implementations use Ollama-hosted embedding models:
- **Llama 3.1 pipeline**: Uses `OllamaEmbeddings` from LangChain with `model="llama3.1"`
- **Gemma pipeline**: Uses `OllamaEmbedder` from Agno with `id="embeddinggemma:latest"`

These models run on your local Ollama server at `http://127.0.0.1:11434`, converting text chunks into vector representations without network transmission.

### Vector Storage Options

The repository demonstrates two lightweight local vector databases:
- **Chroma**: In-memory storage ideal for session-based RAG with Llama 3.1
- **LanceDB**: File-based persistent storage used with Gemma embeddings, allowing knowledge bases to survive across application restarts

### Local LLM Inference

Generation stays local using Ollama-hosted chat models:
- `ChatOllama` (LangChain) for the Llama 3.1 implementation
- `Ollama` (Agno) for the Gemma-based agentic RAG

Both interfaces stream tokens directly from your local GPU or CPU, ensuring prompts and responses never leave the machine.

## Implementation 1: Llama 3.1 with LangChain and Chroma

The first implementation in [`rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py) demonstrates a straightforward RAG pipeline for querying web pages and documents using Meta's Llama 3.1 model.

```python

# File: rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py

# https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py

import streamlit as st
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import WebBaseLoader
from langchain_community.vectorstores import Chroma
from langchain_ollama import OllamaEmbeddings, ChatOllama

# 1. UI – ask for a webpage URL

webpage_url = st.text_input("Enter Webpage URL")

if webpage_url:
    # 2. Load & split the page

    docs = WebBaseLoader(webpage_url).load()
    splits = RecursiveCharacterTextSplitter(
        chunk_size=500,
        chunk_overlap=10
    ).split_documents(docs)

    # 3. Create local embeddings + vector store

    embeddings = OllamaEmbeddings(
        model="llama3.1",
        base_url="http://127.0.0.1:11434"
    )
    vectorstore = Chroma.from_documents(splits, embeddings)

    # 4. LLM (local Ollama)

    ollama = ChatOllama(
        model="llama3.1",
        base_url="http://127.0.0.1:11434"
    )

    def answer(question):
        # Retrieve most-relevant chunks

        docs = vectorstore.as_retriever().invoke(question)
        context = "\n\n".join(d.page_content for d in docs)

        # Prompt the model (local)

        prompt = f"Question: {question}\n\nContext: {context}"
        response = ollama.invoke([('human', prompt)])
        return response.content.strip()

    # 5. Ask a question about the page

    query = st.text_input("Ask any question about the webpage")
    if query:
        st.write(answer(query))

```

**Key implementation details:**

- **Local-only processing**: The `base_url="http://127.0.0.1:11434"` parameter ensures all embedding and generation requests route to your local Ollama instance.
- **In-memory vector store**: `Chroma.from_documents()` creates a temporary vector database that exists only for the session, perfect for ephemeral document analysis.
- **Simple retrieval**: The `as_retriever().invoke()` method performs semantic search against the local Chroma store, returning the top-k relevant chunks.

## Implementation 2: Gemma Embeddings with Agno and LanceDB

The second implementation in [`rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py) showcases an agentic approach using Google's Gemma for embeddings and LanceDB for persistent storage.

```python

# File: rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py

# https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py

import streamlit as st
from agno.agent import Agent
from agno.knowledge.embedder.ollama import OllamaEmbedder
from agno.knowledge.knowledge import Knowledge
from agno.models.ollama import Ollama
from agno.vectordb.lancedb import LanceDb, SearchType

# 1. Set up a persistent knowledge base that uses the Gemma embedder

@st.cache_resource
def load_kb():
    return Knowledge(
        vector_db=LanceDb(
            table_name="docs",
            uri="tmp/lancedb",
            search_type=SearchType.vector,
            embedder=OllamaEmbedder(
                id="embeddinggemma:latest", 
                dimensions=768
            ),
        )
    )

kb = load_kb()

# 2. UI – let the user add PDFs/URLs

new_url = st.sidebar.text_input(
    "Add URL", 
    placeholder="https://example.com/file.pdf"
)
if st.sidebar.button("Add"):
    if new_url and new_url not in st.session_state.get("urls", []):
        st.session_state.setdefault("urls", []).append(new_url)
        kb.add_content(url=new_url)  # Gemma creates the embeddings locally

# 3. Build an Agno Agent that uses a local Llama 3.2 model for generation

agent = Agent(
    model=Ollama(id="llama3.2:latest"),
    knowledge=kb,
    instructions=[
        "Search the knowledge base for relevant info and answer based on it.",
        "Use clear headings, bullet points, and concise language."
    ],
    search_knowledge=True,
    markdown=True,
)

# 4. Ask a question – stream the answer back to the UI

question = st.text_input("Enter your question:")
if st.button("Get Answer") and question:
    with st.spinner("Generating answer…"):
        answer = ""
        placeholder = st.empty()
        for chunk in agent.run(question, stream=True):
            answer += chunk.content or ""
            placeholder.markdown(answer)

```

**Key implementation details:**

- **Persistent vector storage**: `LanceDb` with `uri="tmp/lancedb"` writes embeddings to disk, enabling the knowledge base to persist across application restarts without re-indexing.
- **Gemma embeddings**: The `OllamaEmbedder` with `id="embeddinggemma:latest"` generates 768-dimensional vectors locally, providing high-quality semantic representations without cloud dependencies.
- **Agentic orchestration**: The Agno `Agent` class handles retrieval and generation automatically when `search_knowledge=True`, streaming responses token-by-token for real-time UI updates.

## Running the Applications Locally

To deploy these privacy-preserving RAG systems on your own hardware, follow these setup steps:

1. **Install Ollama** from [ollama.com](https://ollama.com) and verify the service is running on `http://127.0.0.1:11434`.

2. **Pull the required models**:

   ```bash
   ollama pull llama3.1
   ollama pull embeddinggemma:latest
   ollama pull llama3.2:latest
   ```

3. **Clone the repository and install dependencies**:

   ```bash
   git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git
   cd awesome-llm-apps
   pip install -r requirements.txt
   ```

4. **Launch the Streamlit applications**:

   ```bash
   # For Llama 3.1 RAG

   streamlit run rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py
   
   # For Gemma embedding RAG

   streamlit run rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py
   ```

5. **Interact via browser** – input URLs or PDFs, then query the knowledge base. All embeddings, retrievals, and generations execute locally on your CPU or GPU.

## Summary

Implementing RAG with local open-source LLMs enables organizations to leverage AI-powered document analysis while maintaining complete data sovereignty. The key takeaways from the awesome-llm-apps implementations include:

- **Zero external dependencies**: Both the Llama 3.1 and Gemma pipelines run entirely through Ollama on `http://127.0.0.1:11434`, ensuring no data leaves your local network.
- **Flexible vector storage**: Choose **Chroma** for ephemeral, in-memory sessions or **LanceDB** for persistent, file-based knowledge bases that survive application restarts.
- **Modular architecture**: Swapping models requires only changing the `model` parameter in `OllamaEmbeddings` or `OllamaEmbedder`, allowing easy experimentation with new open-source releases.
- **Agentic capabilities**: The Agno framework demonstrates how local RAG can evolve into autonomous agents that reason over private knowledge bases without external API costs.

## Frequently Asked Questions

### Can I use these local RAG implementations for commercial documents without violating data privacy regulations?

Yes. Because all processing—including embedding generation, vector storage, and LLM inference—occurs locally via Ollama, your documents never transmit to third-party servers. This architecture supports compliance with GDPR, HIPAA, and SOC 2 requirements that mandate data residency, though you should still verify that your local hardware and Ollama installation meet your organization's specific security standards.

### How do I switch from Llama 3.1 to a different local model like Mistral or Qwen?

Changing the underlying LLM requires only updating the model identifier in your code. For the LangChain implementation in [`rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/llama3.1_local_rag/llama3.1_local_rag.py), change the `model` parameter in both `OllamaEmbeddings` and `ChatOllama` to your preferred Ollama model (e.g., `"mistral"` or `"qwen2.5"`). For the Agno implementation, update the `Ollama` model ID and ensure you have pulled the corresponding model via `ollama pull <model_name>`.

### What are the hardware requirements for running these privacy-preserving RAG systems?

The implementations run on consumer hardware, though performance varies by model size. Llama 3.1 (8B parameters) and Gemma embedding models operate comfortably on machines with 16GB RAM and a modern CPU, though a GPU with 8GB+ VRAM significantly improves inference speed. LanceDB requires minimal disk space (megabytes to gigabytes depending on corpus size), while Chroma operates entirely in memory, requiring sufficient RAM to hold your document chunks and vectors simultaneously.

### Can I extend these implementations to handle images or multi-modal documents?

Yes. The Gemma embedding model (`embeddinggemma:latest`) supports multi-modal inputs, and the Agno framework's architecture in [`rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py`](https://github.com/Shubhamsaboo/awesome-llm-apps/blob/main/rag_tutorials/agentic_rag_embedding_gemma/agentic_rag_embeddinggemma.py) can be extended to process images. You would replace the text-only loaders with libraries like `PyMuPDF` for PDF image extraction or `Pillow` for direct image loading, then pass the processed image data to the embedding model. The LanceDB vector store can store multi-modal embeddings alongside text vectors, enabling RAG over mixed document types while maintaining the same local privacy guarantees.