How to Perform Single-Turn Inference with `agent.complete()` in Needle
Call Needle.complete(prompt, max_new_tokens) to generate text in a single request without function calling, which returns a dict containing the generated content under the "content" key.
In the cactus-compute/needle library, single-turn inference is the simplest way to generate text from a language model. This guide explains the complete execution flow, from the public API through the native C extension, with working code examples you can run immediately.
What Needle.complete() Does Under the Hood
When you invoke complete(), the library executes a five-step pipeline. Understanding this flow helps debug issues and optimize performance.
Step 1: Telemetry Instrumentation
Every call is tracked via _track("complete", …) in needle/__init__.py (line 57). This records usage metrics without blocking the inference request.
Step 2: Dispatch to Internal Implementation
The public complete() method immediately forwards to the private _complete() method (lines 57-60). This indirection allows the public API to remain stable while internal implementation evolves.
Step 3: Engine Binding (Lazy Initialization)
_complete() calls _bind() (lines 24-33) to initialize the native Needle engine:
- Loads the appropriate native library for your hardware generation (2 or 3)
- Executes
needle_initwith your system prompt and tool schema - Skips binding if a worker process is already active (e.g., for fine-tuned weights)
This lazy binding means the first inference call has slightly higher latency than subsequent calls.
Step 4: Inference Execution
Two code paths exist depending on configuration:
Fine-tuned path — When weights are supplied to the constructor, FineTuneWorker spawns a child process. The request is serialized to JSON-RPC and sent to needle/_worker.py (lines 8-13). The worker loads your .cact archive and runs inference in isolation.
Standard engine path — The C-extension function needle_complete is invoked directly. The prompt is UTF-8 encoded, max_new_tokens is passed as an integer, and a pre-allocated ctypes buffer receives the JSON response. Any error code < 0 raises RuntimeError (lines 66-72).
Step 5: Response Parsing
Raw bytes are decoded to UTF-8 and parsed with json.loads (lines 73-81). Parsing failures raise descriptive RuntimeError messages. For fine-tuned models, the confidence field is forced to None — the fine-tuning process does not update the confidence head.
Response Format
complete() returns a Python dict with this structure:
{
"type": "text",
"content": "...generated text..."
}
When tools are defined, additional keys like function_calls appear. For single-turn inference without tools, you only need response["content"].
Code Examples
Basic Single-Turn Inference
from needle import Needle
# Create agent with defaults (no tools, empty system prompt)
agent = Needle()
response = agent.complete(
text="Explain quantum computing in one sentence.",
max_new_tokens=64,
)
print(response["content"])
This triggers the standard engine path described above, ultimately calling the native needle_complete function.
With Custom System Prompt
system_prompt = "You are a concise technical writer who avoids jargon."
agent = Needle(system=system_prompt)
result = agent.complete(
text="What is a neural network?",
max_new_tokens=48,
)
print(result["content"])
The system prompt is passed to needle_init during binding and persists across multiple complete() calls.
With Fine-Tuned Weights
# Path to your .cact archive from Needle's fine-tuning pipeline
agent = Needle(weights="./my_model.cact")
answer = agent.complete(
text="Generate a product description for noise-canceling headphones.",
max_new_tokens=128,
)
print(answer["content"])
The weights parameter causes Needle to instantiate FineTuneWorker from needle/_worker.py. Requests route through a child process that loads your custom weights.
Key Implementation Files
| File | Purpose |
|---|---|
needle/__init__.py |
Public Needle class, complete() and run() methods, engine binding logic |
needle/_worker.py |
FineTuneWorker class for child-process inference with custom weights |
needle/agent/fetch.py |
Native library download, caching, and path resolution |
needle/agent/tools.py |
Tool schema utilities for function-calling mode |
needle/playground/server.py |
Reference HTTP server showing JSON-RPC patterns |
Handling Errors
Two primary failure modes exist:
- Engine errors — Negative return codes from
needle_completeraiseRuntimeErrorwith context from the C extension - JSON parse errors — Malformed responses raise
RuntimeErrorafter decoding fails; this typically indicates version mismatches between the Python package and native library
Always wrap production calls in try/except blocks and log the full error message for debugging.
Performance Considerations
- First-call latency —
complete()executes_bind()on first use, which downloads the native library if missing - Fine-tuned overhead — Child process spawning adds 1-3 seconds for the first request with custom weights
- Buffer reuse — The ctypes buffer is pre-allocated based on
max_new_tokensto minimize memory churn
Summary
Needle.complete(prompt, max_new_tokens)is the entry point for single-turn text generation- Lazy binding initializes the native engine on first call via
_bind()inneedle/__init__.py - Two execution paths: direct C extension for standard inference,
FineTuneWorkerprocess for custom weights - Response is always a dict with
"type"and"content"keys; extractresponse["content"]for generated text - Source files
needle/__init__.pyandneedle/_worker.pyimplement the complete flow
Frequently Asked Questions
How do I set the maximum response length?
Pass max_new_tokens as an integer to complete(). The native engine allocates a proportional buffer and stops generation at that token limit. Higher values increase memory usage but do not guarantee longer outputs if the model reaches a natural stopping point.
What hardware generations does Needle support?
The _bind() method in needle/__init__.py automatically detects and loads generation 2 or 3 libraries. You do not specify this manually; fetch_library in needle/agent/fetch.py selects the correct binary for your system.
Why is my fine-tuned model missing confidence scores?
The FineTuneWorker implementation forces confidence: None in the response. According to the source code in needle/__init__.py (lines 78-81), the fine-tuning pipeline does not update the confidence head, so this field is intentionally disabled to avoid misleading values.
Can I use complete() with function calling tools?
No — single-turn complete() is designed for text generation only. For tool use, call agent.run() which handles the multi-turn conversation loop required for function execution. Tool schemas are still passed to needle_init during binding, but complete() ignores any function_calls in responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →