RAM Requirement for Needle 2: 28 MB for Full Inference
Needle 2 requires approximately 28 MB of RAM to run a full inference session while shipping as a single 14 MB binary.
The RAM requirement for Needle 2 is remarkably low compared to conventional language models. According to the cactus-compute/needle source code, the model is packaged as a single 14 MB binary and consumes roughly 28 MB of RAM during active inference. This efficiency stems from the Simple Attention Network architecture, CQ2-bit quantization, and a 256-token sliding window managed by the inference engine.
How Needle 2 Achieves a 28 MB RAM Footprint
The README.md in the cactus-compute/needle repository states that Needle 2 is designed to be extremely lightweight. When running a full inference session, the model consumes about 28 MB of RAM. This low memory footprint is the result of three core optimizations implemented across the codebase.
Simple Attention Network Architecture
In needle/model/architecture.py, the Simple Attention Network architecture replaces traditional transformer blocks with a lighter design. By reducing the parameter count and simplifying attention computations, the architecture keeps both disk size and resident memory to a minimum.
CQ2-Bit Quantization
The file needle/model/quantize.py implements CQ2-bit quantization (Cactus Quants). This aggressive quantization scheme compresses weights to two bits per parameter, directly shrinking the model binary to 14 MB and reducing the memory needed to hold weights during inference.
256-Token Sliding Window Engine
In needle/model/run.py, the NeedleEngine class manages inference by maintaining only a 256-token sliding window in active memory. Instead of loading a full context buffer, the engine discards older tokens beyond this window, which caps the RAM requirement for Needle 2 at approximately 28 MB even during a complete session.
Running Needle 2 Within 28 MB of RAM
All public APIs documented in doc/apis.md are designed to operate within this 28 MB budget. Below are minimal, runnable examples that execute comfortably on memory-constrained hardware.
Tool-Calling Example
import needle
@needle.tool
def get_weather(city: str):
"""Get the current weather for a city."""
# Dummy implementation – replace with a real API call
return {"city": city, "temp_c": 27, "sky": "clear"}
agent = needle.Needle(tools=[get_weather])
result = agent.run("What's the weather like in Lagos right now?")
print(result["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]
Structured Data Extraction Example
from pydantic import BaseModel
import needle
class WeatherInfo(BaseModel):
city: str
temperature: float
condition: str
# Describe the extraction schema via Pydantic
weather_schema = WeatherInfo
text = "In Paris it is 15 °C and raining."
extracted = needle.extract(text, schema=weather_schema)
print(extracted)
# WeatherInfo(city='Paris', temperature=15.0, condition='raining')
Direct Inference with NeedleEngine
from needle.model.run import NeedleEngine
engine = NeedleEngine(model_path="cactus-compute/needle2")
prompt = "Summarize the key points of the meeting."
output = engine.generate(prompt, max_new_tokens=128)
print(output)
Each example above can run on a machine with ≈28 MB of RAM, making Needle 2 suitable for edge devices, containers, or any environment where memory is scarce.
Summary
- The RAM requirement for Needle 2 is approximately 28 MB for a full inference session.
- The model itself is distributed as a single 14 MB binary.
needle/model/architecture.pydefines the Simple Attention Network that reduces parameter overhead.needle/model/quantize.pyapplies CQ2-bit quantization to compress model weights.needle/model/run.pyprovides theNeedleEnginethat enforces a 256-token sliding window, keeping active memory low.
Frequently Asked Questions
What is the RAM requirement for Needle 2?
Needle 2 requires approximately 28 MB of RAM to execute a full inference session. This figure is documented in the repository's README.md and is maintained by the 256-token sliding window managed in needle/model/run.py. The lightweight design makes it one of the most memory-efficient inference options available.
How large is the Needle 2 model binary?
The Needle 2 model is packaged as a single 14 MB binary. This small footprint is achieved through CQ2-bit quantization implemented in needle/model/quantize.py, which compresses the model weights before distribution. Despite the tiny binary size, the model remains capable of tool calling and structured extraction.
What keeps Needle 2's memory usage so low?
Three core optimizations work together: the Simple Attention Network defined in needle/model/architecture.py, CQ2-bit quantization in needle/model/quantize.py, and the 256-token sliding window enforced by NeedleEngine in needle/model/run.py. Together these limit both weight storage and active context memory during inference.
Can Needle 2 run on edge devices and in containers?
Yes, the 28 MB RAM footprint and 14 MB binary size make Needle 2 ideal for edge devices, IoT hardware, and containerized environments where memory quotas are tight. All high-level APIs shown in doc/apis.md and the examples above are tested to run within this constrained memory budget.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →