How to Build a RAG Pipeline Using Agno's Knowledge Module with Vector Databases
Agno's Knowledge module provides a unified façade that orchestrates document ingestion, embedding, and vector storage, enabling you to build production-ready RAG pipelines by simply passing a Knowledge instance to an Agent.
The agno-agi/agno library offers a modular Knowledge subsystem that abstracts the complexity of Retrieval-Augmented Generation (RAG) into composable components. By leveraging the Knowledge class alongside pluggable vector database adapters, you can build a RAG pipeline using Agno's Knowledge module with vector databases like PostgreSQL, LightRAG, and Qdrant without rewriting application logic.
Understanding the RAG Architecture
Agno structures its RAG stack into five distinct layers orchestrated by the Knowledge class in libs/agno/agno/knowledge/knowledge.py. Each layer handles a specific concern:
- Data Ingestion: Raw content loading and chunking via
libs/agno/agno/knowledge/reader/*(e.g.,markdown_reader.py,pdf_reader.py) - Embedding: Dense vector generation through
libs/agno/agno/knowledge/embedder/base.pyand concrete providers likeopenai.py - Vector Storage: Persistence and similarity search via
libs/agno/agno/vectordb/*(e.g.,pgvector/pgvector.py,lightrag/lightrag.py) - Reranking (optional): Result refinement using LLM-based rerankers in
libs/agno/agno/knowledge/reranker/* - Knowledge-aware Agent: Automatic tool injection in
libs/agno/agno/agent/agent.pywhen aKnowledgeinstance is supplied
This architecture allows you to swap vector database backends or embedding models by changing constructor arguments rather than refactoring pipeline logic.
Step-by-Step Pipeline Construction
1. Initialize Your Vector Database
Select a backend that matches your infrastructure. For PostgreSQL with the pgvector extension:
from agno.vectordb.pgvector import PGVectorDB
vector_db = PGVectorDB(
connection_string="postgresql://user:pwd@localhost:5432/agno",
collection_name="rag_documents",
)
2. Configure the Embedder
Choose from OpenAI, Cohere, Sentence-Transformers, or other providers implemented in libs/agno/agno/knowledge/embedder/:
from agno.knowledge.embedder.openai import OpenAIEmbedder
embedder = OpenAIEmbedder(model="text-embedding-3-large")
3. Instantiate the Knowledge Store
Wire together the vector store and embedder. The Knowledge class in knowledge.py acts as the coordinator:
from agno.knowledge.knowledge import Knowledge
knowledge = Knowledge(
vector_store=vector_db,
embedder=embedder,
# Optional: add reranker=InfinityReranker()
)
4. Ingest Documents
Use built-in readers or implement the BaseReader protocol from libs/agno/agno/knowledge/reader/base.py. The upsert_many() method handles chunking, embedding, and storage atomically:
from agno.knowledge.reader.markdown_reader import MarkdownReader
docs = MarkdownReader().load(path="docs/introduction.md")
knowledge.upsert_many(docs)
5. Create a RAG-Enabled Agent
When you pass the knowledge object to an Agent, Agno automatically adds a search_knowledge tool and configures the retrieval logic:
from agno.agent import Agent
agent = Agent(
model="gpt-4o-mini",
knowledge=knowledge, # RAG enabled automatically
)
6. Execute Queries with Retrieval
The agent performs similarity search, optionally reranks results, and injects retrieved chunks as references in the prompt:
response = agent.run("Explain the core concepts of Retrieval-Augmented Generation.")
print(response.message) # Generated answer
print(response.references) # List of KnowledgeDocument citations
Complete Implementation Example
The following script demonstrates a full pipeline using PGVector and OpenAI:
import os
from agno.vectordb.pgvector import PGVectorDB
from agno.knowledge.embedder.openai import OpenAIEmbedder
from agno.knowledge.knowledge import Knowledge
from agno.knowledge.reader.markdown_reader import MarkdownReader
from agno.agent import Agent
# 1️⃣ Vector DB configuration
vector_db = PGVectorDB(
connection_string=os.getenv("PGVECTOR_URL"),
collection_name="rag_demo",
)
# 2️⃣ Embedding model
embedder = OpenAIEmbedder(model="text-embedding-3-large")
# 3️⃣ Knowledge store assembly
knowledge = Knowledge(vector_store=vector_db, embedder=embedder)
# 4️⃣ Document ingestion
md_reader = MarkdownReader()
documents = md_reader.load(path="cookbook/07_knowledge/knowledge_demo.md")
knowledge.upsert_many(documents)
# 5️⃣ Agent instantiation with RAG
agent = Agent(
model="gpt-4o-mini",
knowledge=knowledge,
temperature=0.0,
)
# 6️⃣ Query execution
result = agent.run("Summarize how Agno handles chunking and vector storage.")
print("Answer:", result.message)
print("\nReferences:")
for ref in result.references:
print(f"- {ref.title} ({ref.id})")
This workflow reads the Markdown file, splits content into semantic chunks, generates embeddings via OpenAI, persists vectors to PostgreSQL, and generates cited responses using the retrieved context.
Swapping Vector Database Backends
Agno's design is plug-and-play. To use LightRAG instead of PGVector, change only the vector store instantiation:
from agno.vectordb.lightrag import LightRag
vector_db = LightRag(
api_key=os.getenv("LIGHTRAG_API_KEY"),
server_url=os.getenv("LIGHTRAG_SERVER_URL", "http://localhost:9621"),
collection_name="rag_collection",
)
# Knowledge and Agent configuration remain identical
The Knowledge façade in knowledge.py abstracts backend-specific upload logic, as implemented in libs/agno/agno/vectordb/lightrag/lightrag.py.
Core Source Files Reference
| Component | File Path | Purpose |
|---|---|---|
| Knowledge Façade | libs/agno/agno/knowledge/knowledge.py |
Orchestrates reading, embedding, upsert, search, and optional reranking |
| Base Embedder | libs/agno/agno/knowledge/embedder/base.py |
Abstract interface for all vectorization providers |
| OpenAI Embedder | libs/agno/agno/knowledge/embedder/openai.py |
Calls OpenAI's embedding endpoint |
| PGVector Adapter | libs/agno/agno/vectordb/pgvector/pgvector.py |
Stores vectors using PostgreSQL's pgvector extension |
| LightRAG Adapter | libs/agno/agno/vectordb/lightrag/lightrag.py |
Remote LightRAG server integration |
| Base Reader | libs/agno/agno/knowledge/reader/base.py |
Protocol for custom content loaders |
| Markdown Reader | libs/agno/agno/knowledge/reader/markdown_reader.py |
Parses Markdown into KnowledgeDocument objects |
| RAG Agent | libs/agno/agno/agent/agent.py |
Automatically adds search_knowledge tool when Knowledge is supplied |
| Reranker Base | libs/agno/agno/knowledge/reranker/base.py |
Interface for LLM-based result refinement |
Summary
- The
Knowledgeclass inknowledge.pyserves as the central coordinator for RAG pipelines, abstracting chunking, embedding, and retrieval. - Vector database backends are interchangeable via the
vectordbmodule; swapPGVectorDBforLightRagor others without changing pipeline logic. - Agents automatically enable RAG when instantiated with a
knowledgeparameter, adding thesearch_knowledgetool and injecting retrieved references into prompts. - Document readers in
knowledge/reader/handle parsing, whileupsert_many()manages the embedding and storage workflow. - Source citations are available via
response.references, providing traceability to the original documents.
Frequently Asked Questions
What is the Knowledge class in Agno?
The Knowledge class is a high-level façade defined in libs/agno/agno/knowledge/knowledge.py that orchestrates the entire RAG workflow. It combines a VectorDB instance, an Embedder, and optional Reranker components to handle document ingestion, vectorization, storage, and retrieval through a unified API.
Which vector databases does Agno support?
Agno supports multiple backends through adapters in libs/agno/agno/vectordb/, including PostgreSQL with pgvector (PGVectorDB), LightRAG (LightRag), Qdrant, and Pinecone. You can switch between these by changing the vector_store parameter in the Knowledge constructor without modifying downstream code.
How does document chunking work in the pipeline?
Chunking occurs automatically during the ingestion phase when you call knowledge.upsert_many(). The specific strategy depends on the reader implementation; for example, MarkdownReader splits content based on semantic structure, while other readers may use size-based or delimiter-based chunking defined in libs/agno/agno/knowledge/reader/.
Can I use custom embedding models or local embedders?
Yes. Agno's embedder architecture in libs/agno/agno/knowledge/embedder/base.py allows you to implement custom providers by subclassing the base class. The library includes built-in support for OpenAI, Cohere, and Sentence-Transformers, enabling both cloud-based and local embedding execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →