How to Perform Structured Extraction Using Needle: A Complete Guide
Use the extract function from needle/__init__.py to perform one-shot structured extraction by passing your text and a Pydantic model or dict schema; the function creates a temporary agent, forces the LLM to emit a matching function call, and returns a typed object.
Needle provides a lightweight, purpose-built API for structured extraction that eliminates complex prompt engineering. This guide covers the core extract function, its implementation in the needle codebase, and practical patterns for extracting typed data from unstructured text.
The extract Function: Core API
The primary entry point for structured extraction is extract, defined at lines 79‑92 of needle/__init__.py.
def extract(text: str, schema: type | dict, system: str | None = None,
max_new_tokens: int = 256, weights: str | None = None) -> object:
"""One‑shot structured extraction:
• `schema` is registered as the sole tool.
• The shared engine is (re‑)initialised with this schema.
• Returns a Pydantic model instance when `schema` is a model,
otherwise a plain dict."""
_track("extract", {"n_tools": 1, "tuned": bool(weights or _active_weights)})
agent = Needle(tools=[schema], system=system, weights=weights or _active_weights)
response = agent._complete(text, max_new_tokens)
calls = response.get("function_calls") or []
if not calls:
return None
arguments = calls[0].get("arguments") or {}
return schema(**arguments) if _is_pydantic_model(schema) else arguments
The function signature reveals three key behaviors:
schema– Accepts either a PydanticBaseModelsubclass or a dict mapping field names to typessystem– Optional system prompt to guide the LLM's extraction behaviorweights– Path to custom fine-tuned weights (defaults to bundled weights)
How Structured Extraction Works Internally
The extract function implements a four-step pipeline that guarantees deterministic, schema-compliant output.
Step 1: Single-Tool Agent Creation
extract instantiates a temporary Needle agent with only your schema as a registered tool. This is implemented in needle/agent/tools.py, where the Needle class stores tools and builds runtime descriptors that the LLM can invoke.
agent = Needle(tools=[schema], system=system, weights=weights or _active_weights)
By registering the schema as the sole tool, the LLM has no alternative but to emit a function call matching that structure.
Step 2: Prompt Dispatch via _complete
The agent's _complete method sends your text to the compiled native runtime. As shown in needle/__init__.py, this calls the underlying _lib().needle_complete function:
response = agent._complete(text, max_new_tokens)
Step 3: Function-Call Extraction
The LLM returns a JSON-encoded response containing function calls. The extract function pulls the first call from response.get("function_calls") at lines 88‑90:
calls = response.get("function_calls") or []
if not calls:
return None
arguments = calls[0].get("arguments") or {}
Step 4: Type-Safe Deserialization
Finally, arguments are mapped onto your schema. At lines 91‑92, the code checks whether the schema is a Pydantic model:
return schema(**arguments) if _is_pydantic_model(schema) else arguments
- Pydantic models → Instantiated and validated automatically
- Dict schemas → Raw dict returned for flexibility
Practical Examples
Simple Dict-Based Extraction
For quick prototyping, pass a dict describing expected fields and types:
from needle import extract
schema = {"city": str, "country": str}
result = extract("The capital of France is Paris, France.", schema)
print(result)
# → {'city': 'Paris', 'country': 'France'}
Pydantic Model Extraction (Recommended)
For production code, use Pydantic models to get validation, type checking, and IDE support:
from pydantic import BaseModel
from needle import extract
class WeatherReport(BaseModel):
city: str
temperature_c: float
condition: str
text = "Weather in Berlin: 23.5°C, partly cloudy."
report = extract(text, WeatherReport)
print(report)
# → WeatherReport(city='Berlin', temperature_c=23.5, condition='partly cloudy')
print(report.dict())
# → {'city': 'Berlin', 'temperature_c': 23.5, 'condition': 'partly cloudy'}
Multi-Tool Workflows with Full Agent
When you need extraction alongside other capabilities, use the Needle class directly. This pattern from needle/agent/tools.py lets you combine your schema with custom tools:
from needle import Needle, tool
from pydantic import BaseModel
from typing import Annotated
@tool
def log(message: Annotated[str, "Log message to output"]):
print("[LOG]", message)
class Contact(BaseModel):
name: str
email: str
agent = Needle(tools=[Contact, log])
# The LLM can both extract structured data AND invoke your custom tools
response = agent.run("Send an email to John Doe <john@doe.com>")
The agent.run method returns a response containing both the extracted Contact model and any log calls triggered during processing.
Configuration Options
| Parameter | Default | Purpose |
|---|---|---|
system |
None |
System prompt guiding extraction style |
max_new_tokens |
256 |
Maximum tokens for LLM response |
weights |
None |
Path to custom fine-tuned weights |
Pass weights to use a domain-specific fine-tuned model:
from needle import extract
class MedicalTerm(BaseModel):
term: str
icd10_code: str
result = extract(
"Patient diagnosed with Type 2 diabetes mellitus",
MedicalTerm,
weights="/path/to/medical-tuned-weights.safetensors"
)
Key Source Files
Understanding these files deepens your ability to debug and extend Needle's extraction capabilities:
needle/__init__.py– Public API surface includingextractand telemetry tracking via_trackneedle/agent/tools.py– CoreNeedleclass with tool registration and runtime bindingsneedle/model/run.py– Low-level wrappers for LLM prompt executiontests/test_inference.py– Reference implementations demonstratingContactmodel extraction
Summary
extractinneedle/__init__.pyprovides the fastest path to structured extraction—one function call with your schema and text- Single-tool agent architecture guarantees schema compliance by eliminating alternative function choices
- Pydantic models give you automatic validation, IDE autocomplete, and clean
.dict()serialization Needleclass unlocks multi-tool workflows when you need extraction plus custom actions- Custom weights support domain-specific fine-tuning via the
weightsparameter
Frequently Asked Questions
What happens if the LLM fails to return a function call?
The extract function returns None when response.get("function_calls") is empty. You should handle this case in production code:
result = extract(text, MySchema)
if result is None:
# Fall back to retry with different prompt or raise error
pass
This behavior is implemented at line 89 of needle/__init__.py.
Can I use nested Pydantic models for complex schemas?
Yes. The _is_pydantic_model check and schema(**arguments) instantiation support arbitrarily nested Pydantic models. The LLM receives a JSON Schema representation of your full model hierarchy, and nested structures deserialize automatically.
How does extract differ from using Needle directly?
extract creates a temporary, single-purpose agent with your schema as the only tool, optimized for one-shot extraction. The Needle class supports persistent agents with multiple tools, conversation history, and repeated interactions. Use extract for simple data extraction; use Needle for interactive agents or multi-step workflows.
Does structured extraction work with custom fine-tuned models?
Yes. Pass the path to your weights via the weights parameter in either extract or Needle. The library reinitializes the engine with your weights and tracks fine-tuned usage via _track("extract", {"tuned": True}) as shown at line 84 of needle/__init__.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →