# How to Merge LoRA Adapters into a Base Model and Build a Tuned .cact Binary File

> Learn to merge LoRA adapters into a base model and build a tuned .cact binary file using the Needle repository's streamlined three-step workflow. Optimize your model efficiently.

- Repository: [Cactus Compute, Inc./needle](https://github.com/cactus-compute/needle)
- Tags: how-to-guide
- Published: 2026-08-14

---

**The Needle repository provides a three-step workflow: train a LoRA adapter with `finetune_local`, merge it into base weights using `build_main` and `merge_lora`, then export to a compressed .cact binary with `write_export`.**

The Needle project implements a lightweight, JAX-based inference engine that supports efficient fine-tuning through Low-Rank Adaptation (LoRA). This approach lets you adapt large pretrained models without modifying the original checkpoint, then bake those adaptations into a single portable binary. Below you'll find the complete technical walkthrough derived from the cactus-compute/needle source code.

## What is LoRA in the Needle Framework?

**LoRA (Low-Rank Adaptation)** reduces fine-tuning costs by learning small "A" and "B" matrices for specific weight groups instead of updating full parameter tensors. In Needle, these adapters target attention projections—`q_proj`, `k_proj`, `v_proj`, `gate_proj`, and `out_proj`—with configurable rank and scaling.

The architecture lives in [`needle/model/architecture.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/architecture.py), while the training and merging logic resides in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py).

## Step 1: Train a LoRA Adapter with `finetune_local`

The `finetune_local` function in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py) handles the entire training pipeline. It loads the base checkpoint from `checkpoints/needle2.pkl`, initializes LoRA matrices for the target projections, and runs a JAX training loop against your JSONL dataset.

Key implementation details from lines 67–84 and 94–104:

- Adapter creation logic builds A and B matrices with shape determined by `--lora_rank`
- The `--lora_alpha` parameter controls the scaling factor applied during forward passes
- Output serializes to a pickle file at your specified `--checkpoint_dir`

Here's a complete training script that mirrors the `needle finetune` CLI command:

```python
from needle.model.finetune import finetune_local, argparse

parser = argparse.ArgumentParser()
parser.add_argument("--jsonl_path", required=True)
parser.add_argument("--checkpoint_dir", default="ckpt")
parser.add_argument("--lora_rank", type=int, default=8)
parser.add_argument("--lora_alpha", type=float, default=16)
parser.add_argument("--epochs", type=int, default=3)
parser.add_argument("--batch_size", type=int, default=32)
args = parser.parse_args([
    "--jsonl_path", "needle_data.jsonl",
    "--checkpoint_dir", "ckpt",
    "--lora_rank", "8",
    "--lora_alpha", "16",
    "--epochs", "3",
    "--batch_size", "32",
])

finetune_local(args)

```

This produces `ckpt/needle_lora.pkl` containing the trained adapter weights.

## Step 2: Merge the Adapter with `build_main` and `merge_lora`

Once you have an adapter, the `build_main` function (lines 12–34 of [`finetune.py`](https://github.com/cactus-compute/needle/blob/main/finetune.py)) orchestrates the merge operation. It reconstructs LoRA parameters from the pickle file, applies `merge_lora` to combine low-rank updates with base weights, then prepares the unified parameter set for export.

The merge process (line 18):

1. **Loads base checkpoint** via `load_checkpoint` from [`needle/model/run.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/run.py)
2. **Deserializes adapter** with `pickle.load` (lines 12–18)
3. **Applies `merge_lora`** to compute: `W_merged = W_base + (alpha / rank) * B @ A`

```python
from needle.model.finetune import build_main, argparse

parser = argparse.ArgumentParser()
parser.add_argument("--checkpoint", default="checkpoints/needle2.pkl")
parser.add_argument("--lora", required=True)
parser.add_argument("--out", default="needle_finetuned.cact")
parser.add_argument("--bits", default="4")
args = parser.parse_args([
    "--checkpoint", "checkpoints/needle2.pkl",
    "--lora", "ckpt/needle_lora.pkl",
    "--out", "needle_finetuned.cact",
    "--bits", "4",
])

build_main(args)

```

The merged weights now contain the fine-tuned knowledge directly in the parameter tensors—no adapter inference overhead remains.

## Step 3: Export to .cact Format with `write_export`

The final phase invokes `write_export` from [`needle/model/export.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/export.py) (called at lines 25–31 of `build_main`). This serializes:

- Merged model parameters (optionally quantized)
- Architecture configuration
- Tokenizer vocabulary and encoding tables
- KV-window metadata for efficient caching

The default `--bits 4` applies 4-bit quantization, substantially reducing file size while preserving inference quality. Lines 32–34 of `build_main` report the resulting tensor count and binary size.

Loading the finished binary requires no additional files:

```python
from needle import Needle

model = Needle(weights="needle_finetuned.cact", tools=[...])
result = model("What are the top-3 most recent issues in the repo?")
print(result)

```

The `Needle` class automatically restores tokenizer and KV-cache configuration from the binary header.

## Key Source Files You Should Know

| File | Purpose | Critical Functions |
|------|---------|------------------|
| [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py) | LoRA training and merging | `finetune_local`, `build_main`, `merge_lora` |
| [`needle/model/run.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/run.py) | Checkpoint loading | `load_checkpoint` |
| [`needle/model/export.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/export.py) | Binary serialization | `write_export` |
| [`needle/cli.py`](https://github.com/cactus-compute/needle/blob/main/needle/cli.py) | Command-line interface | Exposes `needle finetune`, `needle build` |
| [`needle/model/architecture.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/architecture.py) | Model definition | `SimpleAttentionNetwork` |
| [`needle/model/tokenizer.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/tokenizer.py) | Text encoding | Tokenizer class |

## How to Choose LoRA Hyperparameters

The `finetune_local` signature exposes two critical parameters:

- **`--lora_rank`** (default 8): Controls the inner dimension of A and B matrices. Higher ranks capture more complex adaptations but increase adapter size and training time.
- **`--lora_alpha`** (default 16): Sets the scaling factor applied to LoRA outputs. The effective scale is `alpha / rank`, so the default yields 2x scaling.

For small datasets or narrow task adaptations, ranks of 4–8 typically suffice. Cross-task generalization may require 16–32.

## Summary

- **Train adapters** with `finetune_local` in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py) to create serialized `.pkl` files without touching base weights
- **Merge adapters** via `build_main`, which calls `merge_lora` to bake low-rank updates into the original parameter tensors
- **Export binaries** through `write_export` in [`needle/model/export.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/export.py) to produce self-contained `.cact` files with optional 4-bit quantization
- **Load directly** into the `Needle` class for inference—no separate tokenizer or configuration files required

## Frequently Asked Questions

### What is the difference between a .pkl adapter and a .cact binary?

A `.pkl` file contains only the small LoRA matrices (A and B tensors) plus metadata—it cannot run inference alone and must be combined with a base checkpoint. A `.cact` file is a complete, standalone artifact with merged weights, tokenizer, and inference configuration ready for immediate use.

### Does merging LoRA adapters change the base checkpoint?

No. The `load_checkpoint` function in [`needle/model/run.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/run.py) loads base weights into memory; `merge_lora` operates on these in-memory tensors. The original `checkpoints/needle2.pkl` remains unmodified. You can merge multiple adapters sequentially or maintain separate .cact variants from one base model.

### How much does quantization affect model quality?

The default 4-bit quantization in `write_export` uses symmetric rounding optimized for transformer attention weights. For most fine-tuning scenarios on reasoning tasks, quality degradation is minimal (<2% on needle's evaluation suite). You can specify `--bits 8` or `--bits 16` for higher precision at larger file sizes.

### Can I merge multiple LoRA adapters into one .cact file?

The current `build_main` implementation accepts a single `--lora` path. To compose multiple adapters, train a single adapter on combined data, or sequentially merge adapters by calling `build_main` with the output of a previous merge as the new `--checkpoint`. The pipeline is designed for single-adapter workflows for simplicity.