How to Export Needle 2 Models: Complete Guide to .cact Format and CLI
Needle 2 models export to a binary .cact archive containing quantized weights, tokenizer, and architecture metadata via the needle build command or the write_export() Python API.
Exporting a Needle 2 model transforms your JAX checkpoint into a single portable file that the C++ inference engine can memory-map directly. This guide covers the export workflow, the W4A8 quantization format, and both CLI and programmatic approaches based on the cactus-compute/needle source code.
Understanding the Needle 2 Export Format
Needle 2 uses a self-describing binary format called .cact (Cactus Archive). The runtime loads these files without additional configuration.
Architecture of a .cact File
The format in needle/model/export.py organizes data in three sections:
| Section | Size/Purpose |
|---|---|
| Header (120 bytes) | 30-field geometry descriptor including vocab size, d-model, head count, and shared Lloyd-Max codebooks for CQ quantizers |
| Tensor directory | Fixed-size records (REC_SIZE) with dtype, shape, offset, and size for every tensor |
| Tensor blobs | Quantized weights, FP16 norms, and raw tokenizer data, 64-byte aligned |
Tensors follow a canonical order: embedding → per-layer weights → engram tables → final norm → optional heads → tokenizer.
Quantization Scheme
The format implements W4A8 by default: int4-quantized weights with FP16 activations. The exporter uses:
_cq_pack— Converts weight matrices to codebook-quantized representations using shared codebooks from_cq_codebook_np_pack_lsb— Packs quantized values LSB-first into bitstreams_pack_ternary_crumbs— Handles ternary quantization when specified_align— Ensures 64-byte boundary alignment for each blob
Export Methods: CLI vs. Python API
Method 1: Command-Line Export (Recommended)
The build subcommand in needle/cli.py (lines 66-73) provides the simplest export path.
Basic Export
needle build path/to/checkpoint.pkl --out needle2.cact
Export with LoRA Merge and 2-Bit Quantization
needle build needle-2-base.pkl \
--lora my_adapter.pkl \
--out needle2_2bit.cact \
--bits 2
Available flags:
--bits—"2"or"4"(defaults to checkpoint'sweight_bitsor 4-bit)--bits-map— Custom per-layer precision, e.g.,"2,4,4,4,4"for mixed precision--lora— Path to LoRA adapter for merging before export--out— Destination.cactfile path
Method 2: Python API for Custom Workflows
For embedding export in training pipelines or custom quantization strategies, use needle/model/export.py directly:
from needle.model.run import load_checkpoint
from needle.model.export import write_export
# Load checkpoint with JAX parameters and TransformerConfig
params, config, _ = load_checkpoint("checkpoint.pkl", return_run=True)
# Export with default 4-bit quantization
info = write_export(
params,
config,
path="my_model.cact",
bits=4,
group=128, # CQ quantization group size
tokenizer=None, # Auto-load matching tokenizer
)
print(f"Exported {info['bytes']/1e6:.2f} MiB, {info['tensors']} tensors")
Mixed-Precision Export Example
from needle.model.run import load_checkpoint
from needle.model.export import write_export
params, config, _ = load_checkpoint("checkpoint.pkl", return_run=True)
# First layer 2-bit, remaining layers 4-bit
bits_map = "2,4,4,4,4"
info = write_export(
params,
config,
path="mixed_precision.cact",
bits_map=bits_map,
group=128,
)
The bits argument passes through parse_bits_map (line 83 of export.py), while kv_window injects automatically (lines 58-60).
How Export Works: Internal Pipeline
When write_export() executes, needle/model/export.py runs this pipeline:
- Geometry extraction (
_geometry, lines 103-119) — Validates single head-dimension, rejects unsupported lexicon or sliding-window features - Tensor construction (
_tensors, lines 40-86) — Walks JAX pytree, applies_qfor CQ quantization or_fp16for raw storage - Head tensors (
_head_tensors, lines 92-107) — Addscontrastive_headorconfidence_headif present in checkpoint - Tokenizer blob (
_tokenizer_blob, lines 17-38) — Serializes SentencePiece tokenizer to raw binary viaparse_tokenizer_blob - Packing (
_pack_cact) — Builds aligned header, directory of struct-packed records (_REC_FMT, line 98), and concatenates blobs - Atomic write — Outputs final byte buffer to disk
The resulting file contains all information needed for inference: no external config files required.
Loading Exported Models for Inference
After export, load the .cact file with the Needle runtime:
from needle import Needle
from needle.model.export import RefTokenizer
# Extract tokenizer embedded in the archive
tokenizer = RefTokenizer.from_cact("needle2.cact")
# Initialize agent with exported weights
agent = Needle(weights="needle2.cact", tools=[my_tool])
# Execute inference
response = agent.run("What is the weather in Paris?")
print(response["results"])
The RefTokenizer.from_cact() method reads the raw tokenizer blob directly from the exported archive without separate file management.
Key Source Files
| File | Purpose |
|---|---|
needle/model/export.py |
Core exporter: write_export(), build_export(), binary packing |
needle/model/run.py |
Checkpoint loading: load_checkpoint() |
needle/model/quantize.py |
CQ quantization: _cq_pack(), _cq_unpack(), codebook generation |
needle/model/architecture.py |
TransformerConfig, effective_kv_window() |
needle/cli.py |
CLI entry point, build subcommand registration |
needle/model/tokenizer.py |
SentencePiece construction and serialization |
Summary
- Needle 2 exports to
.cact— A single binary with quantized weights, tokenizer, and geometry - Two export paths —
needle buildCLI for standard workflows,write_export()API for custom pipelines - Default W4A8 quantization — Configurable to 2-bit or mixed precision via
--bitsorbits_map - LoRA merging supported — Adapter weights merge into base checkpoint during export
- Self-describing format — Runtime loads without external metadata; use
RefTokenizer.from_cact()for inference
Frequently Asked Questions
What quantization options does Needle 2 export support?
Needle 2 supports 2-bit, 4-bit, and mixed-precision quantization through the bits and bits_map parameters. Setting --bits 2 applies uniform 2-bit quantization across all layers. The bits_map string like "2,4,4,4,4" assigns per-layer precision—useful for keeping early layers at higher precision while compressing deeper layers.
How do I merge a LoRA adapter when exporting?
Pass the --lora flag to the CLI: needle build base.pkl --lora adapter.pkl --out merged.cact. This triggers the same merge logic used in needle/model/finetune.py's build_main command, combining adapter weights into base parameters before quantization and serialization.
Why does my exported .cact file include the tokenizer?
The .cact format embeds the tokenizer as a RAW blob in the final section. This design makes models fully self-contained—no separate tokenizer files to manage or version-mismatch issues. The runtime extracts it automatically via RefTokenizer.from_cact() without external dependencies.
Can I export from a training checkpoint without running full training?
Yes. The load_checkpoint() function in needle/model/run.py reads any JAX-serialized .pkl checkpoint, including intermediate training saves. You don't need a completed training run; export works as soon as valid parameters and a TransformerConfig exist in the checkpoint file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →