`agent.run()` vs `agent.complete()` in Needle: Key Differences Explained
agent.complete() performs a single inference step, while agent.run() executes a full reasoning loop with automatic tool invocation and result feeding.
When building agents with the Needle library from cactus-compute, understanding the distinction between these two methods is critical for choosing the right interaction pattern. Both methods live in [needle/__init__.py](https://github.com/cactus-compute/needle/blob/main/needle/__init__.py) but serve fundamentally different purposes in the model execution pipeline.
What agent.complete() Does
complete() is a thin wrapper around a single inference call. It invokes the low-level _complete() method, which directly calls needle_complete from the underlying C library.
The method returns a dictionary containing:
type: Typically"response"or"function_call"answer: The model's textual output when applicablefunction_calls: A list of tool calls the model wants to execute (if any)confidence:Noneor a confidence score when using a tuned.cactcheckpoint (lines 35-37 of_complete)
No Python functions are executed. The engine state remains unchanged beyond the returned buffer.
import needle
agent = needle.Needle(tools=[]) # empty tool set
resp = agent.complete("What is the capital of France?")
print(resp)
# {'type': 'response', 'answer': 'Paris', ...}
Use complete() when you need raw model output without side effects—ideal for one-shot extraction, debugging engine responses, or building custom execution logic.
What agent.run() Does
run() implements the full autonomous agent loop. It repeatedly calls _complete(), executes any requested tools, feeds results back to the model, and continues until the model stops requesting tools or reaches max_steps.
The implementation follows this pattern in [needle/__init__.py](https://github.com/cactus-compute/needle/blob/main/needle/__init__.py):
- Call
_complete()for initial inference - Enter loop:
for _ in range(max_steps): - Extract
function_callsfrom response - Look up Python callable:
self._functions.get(name) - Execute call and capture result
- Feed JSON-encoded results back via another
_complete()invocation - Repeat until no more calls or limit reached
The final response includes an additional "results" key aggregating all tool outputs.
import needle
@needle.tool
def get_weather(city: str):
"""Return a mock weather payload."""
return {"city": city, "temp_c": 23, "sky": "cloudy"}
agent = needle.Needle(tools=[get_weather])
resp = agent.run("Tell me the weather in Berlin.")
print(resp["results"])
# [{'city': 'Berlin', 'temp_c': 23, 'sky': 'cloudy'}]
Side-by-Side Comparison
| Aspect | agent.complete() |
agent.run() |
|---|---|---|
| Inference calls | Single | Multiple (loop until done) |
| Tool execution | None | Automatic with result aggregation |
| Return value | Raw engine response | Response plus "results" list |
| Side effects | None | Depends on tool implementations |
| Use case | One-shot queries, debugging | End-to-end agent workflows |
Confidence Handling in Both Methods
When using a tuned checkpoint, both methods handle confidence identically. Since run() ultimately returns the last response from _complete(), confidence propagation works transparently through the loop.
When to Use Each Method
Choose agent.complete() when:
- You only need the model's textual output
- You're debugging raw engine behavior
- You're implementing custom tool execution logic
- You're using
needle.extractfor structured extraction
Choose agent.run() when:
- You want the model to interact with user-defined tools
- You need automatic tool calling and result integration
- You're building end-to-end agent applications
- You want the full reasoning loop without manual orchestration
Summary
agent.complete()wraps a single inference turn—no tools execute, no loop runsagent.run()orchestrates the complete agent loop with automatic tool invocation and result feeding- Both methods share the same underlying
_complete()implementation in [needle/__init__.py](https://github.com/cactus-compute/needle/blob/main/needle/__init__.py) run()adds a"results"key to aggregate tool outputs across the loop- Use
complete()for raw model access; userun()for autonomous agent behavior
Frequently Asked Questions
Can I call tools manually instead of using agent.run()?
Yes. Call agent.complete(), check for function_calls in the response, execute your tools manually, then call complete() again with the results formatted into the conversation. agent.run() automates this pattern.
Does agent.run() have a step limit?
Yes. The loop runs for _ in range(max_steps) to prevent infinite tool calling. You can configure this parameter when initializing the Needle instance.
Why does agent.run() return None for confidence even with a tuned model?
Confidence is extracted from the final _complete() response. If the checkpoint doesn't include confidence scores or the final turn doesn't generate them, the field remains None. This matches complete() behavior since run() delegates to the same underlying method.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →