How to Build a RAG Pipeline Using Agno's Knowledge Module with Vector Databases

Agno's Knowledge module provides a unified façade that orchestrates document ingestion, embedding, and vector storage, enabling you to build production-ready RAG pipelines by simply passing a Knowledge instance to an Agent.

The agno-agi/agno library offers a modular Knowledge subsystem that abstracts the complexity of Retrieval-Augmented Generation (RAG) into composable components. By leveraging the Knowledge class alongside pluggable vector database adapters, you can build a RAG pipeline using Agno's Knowledge module with vector databases like PostgreSQL, LightRAG, and Qdrant without rewriting application logic.

Understanding the RAG Architecture

Agno structures its RAG stack into five distinct layers orchestrated by the Knowledge class in libs/agno/agno/knowledge/knowledge.py. Each layer handles a specific concern:

This architecture allows you to swap vector database backends or embedding models by changing constructor arguments rather than refactoring pipeline logic.

Step-by-Step Pipeline Construction

1. Initialize Your Vector Database

Select a backend that matches your infrastructure. For PostgreSQL with the pgvector extension:

from agno.vectordb.pgvector import PGVectorDB

vector_db = PGVectorDB(
    connection_string="postgresql://user:pwd@localhost:5432/agno",
    collection_name="rag_documents",
)

2. Configure the Embedder

Choose from OpenAI, Cohere, Sentence-Transformers, or other providers implemented in libs/agno/agno/knowledge/embedder/:

from agno.knowledge.embedder.openai import OpenAIEmbedder

embedder = OpenAIEmbedder(model="text-embedding-3-large")

3. Instantiate the Knowledge Store

Wire together the vector store and embedder. The Knowledge class in knowledge.py acts as the coordinator:

from agno.knowledge.knowledge import Knowledge

knowledge = Knowledge(
    vector_store=vector_db,
    embedder=embedder,
    # Optional: add reranker=InfinityReranker()

)

4. Ingest Documents

Use built-in readers or implement the BaseReader protocol from libs/agno/agno/knowledge/reader/base.py. The upsert_many() method handles chunking, embedding, and storage atomically:

from agno.knowledge.reader.markdown_reader import MarkdownReader

docs = MarkdownReader().load(path="docs/introduction.md")
knowledge.upsert_many(docs)

5. Create a RAG-Enabled Agent

When you pass the knowledge object to an Agent, Agno automatically adds a search_knowledge tool and configures the retrieval logic:

from agno.agent import Agent

agent = Agent(
    model="gpt-4o-mini",
    knowledge=knowledge,  # RAG enabled automatically

)

6. Execute Queries with Retrieval

The agent performs similarity search, optionally reranks results, and injects retrieved chunks as references in the prompt:

response = agent.run("Explain the core concepts of Retrieval-Augmented Generation.")
print(response.message)      # Generated answer

print(response.references)   # List of KnowledgeDocument citations

Complete Implementation Example

The following script demonstrates a full pipeline using PGVector and OpenAI:

import os
from agno.vectordb.pgvector import PGVectorDB
from agno.knowledge.embedder.openai import OpenAIEmbedder
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reader.markdown_reader import MarkdownReader
from agno.agent import Agent

# 1️⃣ Vector DB configuration

vector_db = PGVectorDB(
    connection_string=os.getenv("PGVECTOR_URL"),
    collection_name="rag_demo",
)

# 2️⃣ Embedding model

embedder = OpenAIEmbedder(model="text-embedding-3-large")

# 3️⃣ Knowledge store assembly

knowledge = Knowledge(vector_store=vector_db, embedder=embedder)

# 4️⃣ Document ingestion

md_reader = MarkdownReader()
documents = md_reader.load(path="cookbook/07_knowledge/knowledge_demo.md")
knowledge.upsert_many(documents)

# 5️⃣ Agent instantiation with RAG

agent = Agent(
    model="gpt-4o-mini",
    knowledge=knowledge,
    temperature=0.0,
)

# 6️⃣ Query execution

result = agent.run("Summarize how Agno handles chunking and vector storage.")
print("Answer:", result.message)
print("\nReferences:")
for ref in result.references:
    print(f"- {ref.title} ({ref.id})")

This workflow reads the Markdown file, splits content into semantic chunks, generates embeddings via OpenAI, persists vectors to PostgreSQL, and generates cited responses using the retrieved context.

Swapping Vector Database Backends

Agno's design is plug-and-play. To use LightRAG instead of PGVector, change only the vector store instantiation:

from agno.vectordb.lightrag import LightRag

vector_db = LightRag(
    api_key=os.getenv("LIGHTRAG_API_KEY"),
    server_url=os.getenv("LIGHTRAG_SERVER_URL", "http://localhost:9621"),
    collection_name="rag_collection",
)

# Knowledge and Agent configuration remain identical

The Knowledge façade in knowledge.py abstracts backend-specific upload logic, as implemented in libs/agno/agno/vectordb/lightrag/lightrag.py.

Core Source Files Reference

Component File Path Purpose
Knowledge Façade libs/agno/agno/knowledge/knowledge.py Orchestrates reading, embedding, upsert, search, and optional reranking
Base Embedder libs/agno/agno/knowledge/embedder/base.py Abstract interface for all vectorization providers
OpenAI Embedder libs/agno/agno/knowledge/embedder/openai.py Calls OpenAI's embedding endpoint
PGVector Adapter libs/agno/agno/vectordb/pgvector/pgvector.py Stores vectors using PostgreSQL's pgvector extension
LightRAG Adapter libs/agno/agno/vectordb/lightrag/lightrag.py Remote LightRAG server integration
Base Reader libs/agno/agno/knowledge/reader/base.py Protocol for custom content loaders
Markdown Reader libs/agno/agno/knowledge/reader/markdown_reader.py Parses Markdown into KnowledgeDocument objects
RAG Agent libs/agno/agno/agent/agent.py Automatically adds search_knowledge tool when Knowledge is supplied
Reranker Base libs/agno/agno/knowledge/reranker/base.py Interface for LLM-based result refinement

Summary

  • The Knowledge class in knowledge.py serves as the central coordinator for RAG pipelines, abstracting chunking, embedding, and retrieval.
  • Vector database backends are interchangeable via the vectordb module; swap PGVectorDB for LightRag or others without changing pipeline logic.
  • Agents automatically enable RAG when instantiated with a knowledge parameter, adding the search_knowledge tool and injecting retrieved references into prompts.
  • Document readers in knowledge/reader/ handle parsing, while upsert_many() manages the embedding and storage workflow.
  • Source citations are available via response.references, providing traceability to the original documents.

Frequently Asked Questions

What is the Knowledge class in Agno?

The Knowledge class is a high-level façade defined in libs/agno/agno/knowledge/knowledge.py that orchestrates the entire RAG workflow. It combines a VectorDB instance, an Embedder, and optional Reranker components to handle document ingestion, vectorization, storage, and retrieval through a unified API.

Which vector databases does Agno support?

Agno supports multiple backends through adapters in libs/agno/agno/vectordb/, including PostgreSQL with pgvector (PGVectorDB), LightRAG (LightRag), Qdrant, and Pinecone. You can switch between these by changing the vector_store parameter in the Knowledge constructor without modifying downstream code.

How does document chunking work in the pipeline?

Chunking occurs automatically during the ingestion phase when you call knowledge.upsert_many(). The specific strategy depends on the reader implementation; for example, MarkdownReader splits content based on semantic structure, while other readers may use size-based or delimiter-based chunking defined in libs/agno/agno/knowledge/reader/.

Can I use custom embedding models or local embedders?

Yes. Agno's embedder architecture in libs/agno/agno/knowledge/embedder/base.py allows you to implement custom providers by subclassing the base class. The library includes built-in support for OpenAI, Cohere, and Sentence-Transformers, enabling both cloud-based and local embedding execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →