How to Use Local LLMs with an AI Agent: A Complete Implementation Guide
You can use local LLMs with an AI agent by connecting an Ollama-served model to the Agno framework via the Ollama adapter class, then instantiating an Agent with tools and instructions to perform tasks offline.
The awesome-llm-apps repository demonstrates production-ready patterns for running large language models entirely on local hardware. By combining Ollama for model serving with the Agno agent framework, developers can build sophisticated AI agents that operate without external API dependencies, ensuring data privacy and zero inference costs.
Architecture for Using Local LLMs with AI Agents
The repository implements a three-layer architecture that separates model serving from agent logic. This design allows you to swap local models or migrate to remote providers by changing only the adapter configuration.
Model Adapter Layer
The Model Adapter wraps a locally-served LLM behind a unified Python interface. In starter_ai_agents/ai_travel_agent/local_travel_agent.py, the code imports the Ollama class from agno.models.ollama to handle HTTP communication with the Ollama server running on localhost:11434【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py#L7-L8】.
Agent Orchestration Layer
The Agent Orchestration layer defines an Agent instance that combines the model adapter with system instructions and available tools. The travel agent example configures the agent with a specific local model identifier, role description, and behavioral instructions【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py#L74-L78】.
Tooling and UI Layer
The Tooling & UI layer provides optional capabilities like web search integration or Streamlit interfaces. The travel agent attaches SerpApiTools to enable real-time web search capabilities while maintaining the local LLM for reasoning【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py#L90-L92】.
Setting Up Your Local LLM Environment
Before implementing the code examples, ensure your local environment is configured to serve models.
- Install Ollama and start the server with
ollama serve. - Pull your desired model (e.g.,
ollama pull llama3.2orollama pull gemma:2b). - Install Agno via pip to access the agent framework and Ollama adapter.
The model will be accessible at http://localhost:11434, which the Agno Ollama class uses by default.
Implementation Examples for Local LLM Agents
Minimal Local LLM Agent
This example demonstrates the core pattern for using local LLMs with an AI agent. The code creates an Ollama adapter pointing to a local model, then instantiates an Agent with specific instructions for travel planning.
from agno.agent import Agent
from agno.models.ollama import Ollama
from agno.run.agent import RunOutput
# 1️⃣ Model adapter (points to local Ollama)
local_llm = Ollama(id="llama3.2") # ← uses http://localhost:11434 by default
# 2️⃣ Define the agent
travel_planner = Agent(
name="TravelPlanner",
role="Creates a concise travel itinerary.",
model=local_llm,
description="You are a travel planner that suggests activities for a given destination and number of days.",
instructions=[
"Produce a day‑by‑day itinerary with one activity per day.",
"Never fabricate facts; only use information you have researched."
],
add_datetime_to_context=True,
)
# 3️⃣ Run the agent
prompt = "Rome for 3 days"
output: RunOutput = travel_planner.run(prompt, stream=False)
print(output.content) # → generated itinerary
This implementation is based on the full example in starter_ai_agents/ai_travel_agent/local_travel_agent.py【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py#L74-L98】.
Adding Web Search Tools to Your Local Agent
To enhance a local LLM agent with real-time data, you can attach tools like SerpApiTools. This allows the agent to search the web while processing all reasoning through your local model.
from agno.tools.serpapi import SerpApiTools
# Assume `serp_key` is provided by the user
search_tool = SerpApiTools(api_key=serp_key)
researcher = Agent(
name="Researcher",
role="Searches the web for travel activities.",
model=local_llm,
description="Given a destination, generate search queries and fetch top results.",
instructions=[
"Create 3 search terms based on the destination.",
"For each term, call `search_google`.",
"Return the 10 most relevant URLs."
],
tools=[search_tool],
add_datetime_to_context=True,
)
results = researcher.run("Tokyo 5 days", stream=False)
print(results.content) # → list of URLs
The tool integration pattern appears in starter_ai_agents/ai_travel_agent/local_travel_agent.py【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py#L90-L92】.
Building a Local RAG Agent with Vector Storage
For retrieval-augmented generation (RAG) using local models, combine the Ollama LLM adapter with OllamaEmbedder for local embeddings and connect them to a vector store like Qdrant.
from agno.models.ollama import Ollama
from agno.knowledge.embedder.ollama import OllamaEmbedder
from agno.agent import Agent
# Model and embedder
llm = Ollama(id="llama3.2")
embedder = OllamaEmbedder(id="embeddinggemma:latest", dimensions=768)
rag_agent = Agent(
name="RAGAgent",
role="Answers questions using a local document collection.",
model=llm,
knowledge_base={"embedder": embedder, "vector_store": "qdrant://localhost:6333"},
instructions=[
"Search the vector store for relevant chunks.",
"Compose a concise answer referencing the sources."
],
add_datetime_to_context=True,
)
answer = rag_agent.run("What are the health benefits of walking?", stream=False)
print(answer.content)
This implementation is found in rag_tutorials/local_rag_agent/local_rag_agent.py【/tmp/instagit__5n6_2nn/rag_tutorials/local_rag_agent/local_rag_agent.py#L3-L9】【/tmp/instagit__5n6_2nn/rag_tutorials/local_rag_agent/local_rag_agent.py#L14-L32】.
Key Implementation Files in the Repository
The awesome-llm-apps repository contains several reference implementations that demonstrate how to use local LLMs with AI agents across different use cases:
-
starter_ai_agents/ai_travel_agent/local_travel_agent.py– Complete Streamlit application combining a local Ollama model with SerpAPI search tools for travel planning【/tmp/instagit__5n6_2nn/starter_ai_agents/ai_travel_agent/local_travel_agent.py】. -
rag_tutorials/local_rag_agent/local_rag_agent.py– Shows how to pair a local Ollama LLM with a local vector store (Qdrant) viaOllamaEmbedderfor retrieval-augmented generation【/tmp/instagit__5n6_2nn/rag_tutorials/local_rag_agent/local_rag_agent.py】. -
starter_ai_agents/web_scrapping_ai_agent/local_ai_scrapper.py– Uses a local Ollama model to parse scraped HTML into structured JSON without sending data to external APIs. -
advanced_llm_apps/resume_job_matcher/app.py– FastAPI application demonstrating local LLM inference for resume-job description matching. -
agno/models/ollama.py– Core adapter class that translates Agno agent calls into Ollama HTTP requests, handling streaming and response parsing.
These files illustrate the repeatable pattern: import Ollama, create an Agent, optionally attach tools, then call run() to obtain results—all while keeping the model local and offline.
Summary
To use local LLMs with an AI agent effectively, follow these core principles demonstrated in the awesome-llm-apps repository:
- Use the
Ollamaadapter from the Agno framework to connect to locally-served models viahttp://localhost:11434, enabling complete offline operation. - Separate concerns by isolating the model adapter from agent logic and tooling, making it trivial to swap between local models (e.g.,
llama3.2togemma:2b) or migrate to remote providers. - Augment with tools by attaching utilities like
SerpApiToolsfor web search orOllamaEmbedderwith vector stores for RAG, allowing local agents to access real-time or domain-specific data. - Implement the three-layer pattern: Model Adapter (
Ollamaclass), Agent Orchestration (Agentclass with instructions), and Tooling/UI (Streamlit or FastAPI interfaces).
Frequently Asked Questions
Can I use local LLMs with an AI agent without an internet connection?
Yes, once you have downloaded the model weights via Ollama and started the local server with ollama serve, the Agno-based agent operates entirely offline. The Ollama adapter communicates via localhost:11434, requiring no external API calls during inference. However, if your agent uses tools like SerpApiTools for web search, those specific functions will require internet connectivity while the LLM reasoning remains local.
What is the difference between the Ollama class and the Agent class in the Agno framework?
The Ollama class serves as a model adapter that translates Agno's internal method calls into HTTP requests to your local Ollama server, handling streaming, response parsing, and error management. The Agent class is the orchestration layer that manages conversation state, system instructions, tool execution, and context window management. You instantiate the Ollama adapter first, then pass it as the model parameter when creating an Agent, effectively decoupling the LLM backend from the agent's behavioral logic.
How do I add retrieval-augmented generation (RAG) to a local LLM agent?
To implement RAG with local models, combine the Ollama LLM adapter with OllamaEmbedder for generating local embeddings and connect them to a vector store like Qdrant. In rag_tutorials/local_rag_agent/local_rag_agent.py, the implementation uses OllamaEmbedder(id="embeddinggemma:latest", dimensions=768) to create embeddings locally while using Ollama(id="llama3.2") for text generation【/tmp/instagit__5n6_2nn/rag_tutorials/local_rag_agent/local_rag_agent.py#L3-L9】. The agent's knowledge_base parameter links these components, allowing the agent to retrieve relevant document chunks before generating responses, all without sending data to external APIs.
Which local models work best with the Agno agent framework?
The Agno framework's Ollama adapter supports any model compatible with Ollama, including Llama 3.2, Gemma 2B, Mistral, and CodeLlama. The awesome-llm-apps repository specifically demonstrates llama3.2 for general reasoning tasks and gemma:2b for lightweight operations in the RAG tutorial. When selecting a model, consider your hardware constraints—smaller models like Gemma 2B run efficiently on CPU-only machines, while larger models like Llama 3.2 require GPU acceleration for acceptable latency in agent workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →