How `agent.run()` Facilitates a Full Agentic Loop with Tool Execution in Needle
The agent.run() method implements a complete reason-then-act cycle that iteratively calls the LLM, executes Python tools based on function-call responses, and feeds results back until no more calls remain or a step limit is reached.
The Needle class from the cactus-compute/needle repository provides a lightweight, self-contained agent framework. Its run() method—which users access as agent.run()—orchestrates the full agentic loop that powers autonomous tool use without external orchestration services. This article breaks down how the loop works, where the key logic lives, and how to use it effectively.
The Four Stages of the Agentic Loop
The implementation in needle/__init__.py follows a clear ReAct-style pattern: reason, act, observe, repeat.
Stage 1: Prepare the Initial Request
The run begins by optionally processing audio input and sending the user's prompt to the underlying LLM.
# From needle/__init__.py, lines 73-79
audio = self._prepare_audio(audio) if audio else None
response = self._complete(prompt=query, audio=audio, ...)
The _complete method (defined in needle/_worker.py) returns a dictionary that may contain a "function_calls" entry if the model decides to invoke tools.
Stage 2: Iterate Up to max_steps
A for loop drives the core interaction, bounded by the max_steps parameter:
# Lines 80-83: termination check
for _ in range(max_steps):
if response.get("type") != "call" or not response.get("function_calls"):
break
At each iteration, the agent checks whether the response type is "call" and whether any function calls are present. If neither condition holds, the loop exits naturally—this is how the agent knows it has finished reasoning.
Stage 3: Resolve, Execute, and Handle Errors
When function calls exist, the agent locates and invokes the corresponding Python callables:
- Tool lookup (lines 86-89): Each call name resolves to a Python callable stored in
self._functions, populated from user-supplied tool objects defined inneedle/agent/tools.py. - Execution (lines 90-94): The callable runs with arguments extracted from the LLM response. Errors are caught and converted to
{"error": ...}objects so the agent continues rather than crashing.
# Simplified illustration of the execution logic
fn = self._functions[call_name] # resolve name to callable
result = fn(**call_args) # execute with LLM-provided arguments
results.append(result) # collect for feedback
Stage 4: Feed Results Back to the Model
After executing all calls in a step, the results are JSON-encoded and returned to the model:
# Lines 95-96: feedback loop
response = self._complete(prompt=query, context=json.dumps(results), ...)
This closes the loop—the LLM receives tool outputs, incorporates them into its reasoning, and either emits another function call or generates a final answer.
Stage 5: Package and Return Full History
All intermediate results accumulate in the executed list. Upon loop completion, this history attaches to the final response:
# Lines 97-98: final packaging
response["results"] = executed
return response
Complete Working Example
Here's a runnable pattern for custom tool integration:
from needle import Needle
from needle.agent import tools
def fetch_title(url: str) -> dict:
"""Fetch and return the title of a webpage."""
import requests, bs4
r = requests.get(url, timeout=10)
soup = bs4.BeautifulSoup(r.text, "html.parser")
return {"title": soup.title.string.strip() if soup.title else ""}
# Register the custom tool
agent = Needle(tools=[tools.Tool(name="fetch_title", fn=fetch_title)])
# Execute with controlled iteration depth
result = agent.run(
query="What is the title of https://example.com ?",
max_steps=4, # hard limit on tool interactions
max_new_tokens=128
)
print(result["results"])
# → [{'title': 'Example Domain'}]
Built-In Tools and Default Behavior
Needle ships with pre configured tools in needle/agent/tools.py, including webfetch for page retrieval. Using defaults requires no manual tool registration:
agent = Needle() # loads webfetch, search, and other built-ins
resp = agent.run(
query="Summarize https://github.com/cactus-compute/needle",
max_steps=3
)
print(resp["results"])
# → [{'content': '...fetched and summarized content...'}]
Key Implementation Files
| File | Purpose | Location |
|---|---|---|
needle/__init__.py |
Core Needle class, run() method, agentic loop logic |
[needle/__init__.py](https://github.com/cactus-compute/needle/blob/main/needle/__init__.py) |
needle/agent/tools.py |
Tool helper class, built-in tool definitions (webfetch, search) |
[needle/agent/tools.py](https://github.com/cactus-compute/needle/blob/main/needle/agent/tools.py) |
needle/_worker.py |
Native inference bindings: _complete, _bind, low-level model interaction |
[needle/_worker.py](https://github.com/cactus-compute/needle/blob/main/needle/_worker.py) |
tests/test_run.py |
Test coverage for agent loop and tool execution paths | [tests/test_run.py](https://github.com/cactus-compute/needle/blob/main/tests/test_run.py) |
Summary
agent.run()binds together LLM inference, tool dispatch, and feedback in a single method.- The loop runs up to
max_steps, exiting early when no function calls remain. - Tool resolution happens via
self._functions, populated from objects inneedle/agent/tools.py. - Error resilience is built in—tool failures become structured error objects, not exceptions.
- Full execution history returns under the
"results"key for inspection and debugging.
This design keeps the agent self-contained, deterministic, and transparent about its reasoning trajectory.
Frequently Asked Questions
What happens if a tool raises an exception?
The execution wrapper in needle/__init__.py catches exceptions and converts them to {"error": <message>} dictionaries. These feed back to the model as context, allowing the agent to reason about failures and potentially retry or select alternative tools.
How does the agent know when to stop calling tools?
Termination occurs when the LLM response type is not "call" or when the function_calls list is empty. This check runs at the start of each loop iteration (lines 80-83). The loop also hard-stops after max_steps iterations regardless of model output.
Can I use agent.run() without any tools?
Yes. If no tools are registered or the model never emits a function call, the method returns after the initial _complete call with an empty "results" list. The prompt is still processed and any generated text appears in the response.
Where does the actual LLM inference happen?
The _complete method in needle/_worker.py handles native model inference. This abstraction lets Needle.run() focus on orchestration while _worker.py manages tensor operations, batching, and hardware acceleration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →