# How to Synthesize Training Data Using OpenRouter for Needle

> Easily synthesize training data for Needle using the OpenRouter LLM API. Learn to generate realistic tool-calling examples with the generate-data CLI or Python functions.

- Repository: [Cactus Compute, Inc./needle](https://github.com/cactus-compute/needle)
- Tags: how-to-guide
- Published: 2026-08-16

---

**Needle can automatically generate realistic tool-calling training examples by leveraging the OpenRouter LLM API through the `generate-data` CLI command or Python functions in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py).**

The `cactus-compute/needle` repository provides a built-in pipeline for synthesizing training data without manually writing examples. This guide explains how the OpenRouter integration works, how to use the CLI and Python API, and how to customize the generation process for your specific tool schemas.

## Understanding the Data Synthesis Pipeline

The training data generation system lives in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py). The pipeline combines a structured system prompt, JSON templating, and concurrent API calls to OpenRouter to produce diverse, valid training examples at scale.

### Core Components

The workflow relies on several key components working together:

- **`_GEN_TEMPLATE`** – A JSON template constant (lines 31-45) that defines the shape of each generated example, ensuring consistent output structure across all API responses.
- **`_openrouter`** – Internal function (lines 56-66) that builds the request payload and POSTs to the OpenRouter chat completions endpoint.
- **`generate_examples`** – Batch generator that formats prompts and parses LLM responses into structured examples (lines 80-92).
- **`generate_dataset`** – High-level orchestrator using `ThreadPoolExecutor` for concurrent generation with deduplication (lines 99-145).

## Prerequisites and Configuration

Before generating data, you need a valid OpenRouter API key. The code reads the endpoint URL from the `OPENROUTER_URL` environment variable, defaulting to `https://openrouter.ai/api/v1/chat/completions`.

```bash
export OPENROUTER_API_KEY=sk_your_key_here

# Optional: override the default endpoint

export OPENROUTER_URL=https://openrouter.ai/api/v1/chat/completions

```

No additional configuration files are required—the CLI and Python API handle parameter passing directly.

## Generating Data via the CLI

The `generate-data` sub-command in [`needle/cli.py`](https://github.com/cactus-compute/needle/blob/main/needle/cli.py) (lines 41-50) provides the simplest entry point for synthesizing training data using OpenRouter for Needle.

### Fresh Dataset Generation

Create a new dataset from tool schema definitions:

```bash
needle generate-data \
  --tools path/to/tool_schemas.json \
  --num-samples 500 \
  --batch-size 25 \
  --workers 8 \
  --output generated_data.jsonl

```

**Parameter breakdown:**

- **`--tools`** – Path to a JSON file containing your tool schema definitions.
- **`--num-samples`** – Total examples to generate.
- **`--batch-size`** – Samples per API call (default varies; check `generate_main` in [`finetune.py`](https://github.com/cactus-compute/needle/blob/main/finetune.py) lines 74-91).
- **`--workers`** – Concurrent threads for parallel API requests.
- **`--output`** – Destination JSONL file path.

### Augmenting Existing Data

The `augment_jsonl` function (lines 158-171) enables expanding an existing dataset while preserving original samples:

```bash
needle generate-data \
  --augment seed_data.jsonl \
  --num-samples 200 \
  --output seed_data_augmented.jsonl

```

When using `--augment`, Needle automatically extracts tool schemas from the seed file via `_collect_tools`, then generates additional examples using those same definitions.

## Using the Python API Directly

For programmatic control, import directly from `needle.model.finetune`:

```python
from needle.model.finetune import generate_dataset, augment_jsonl
import json

# Load tool schemas

with open("tools.json") as f:
    tools = json.load(f)

# Generate 300 synthetic examples

samples = generate_dataset(tools, num_samples=300)

```

The returned `samples` is a list of dictionaries containing generated queries and structured answers, with original tool schemas attached for downstream reuse.

### Full API Signature

```python
samples = generate_dataset(
    tools,                    # List[dict]: Tool schema definitions

    num_samples=1000,         # int: Total examples to generate

    model="openai/gpt-4",     # str: OpenRouter model identifier

    temperature=0.8,          # float: Sampling temperature (0-2)

    batch_size=20,            # int: Examples per API call

    workers=10,               # int: Concurrent threads

    max_tokens=2048,          # int: Max response length

)

```

**Key parameters for customization:**

- **`model`** – Any OpenRouter-compatible identifier (e.g., `mistralai/mistral-7b-instruct`, `anthropic/claude-3-opus`, `google/gemma-7b-it`).
- **`temperature`** – Lower values (0.0-0.3) produce more deterministic outputs; higher values (0.7-1.0) increase diversity.
- **`workers`** – Increase for faster generation (respecting OpenRouter rate limits).

### Augmenting via Python

```python

# Add 150 samples to existing.jsonl, write to new file

augmented_path = augment_jsonl(
    "existing.jsonl",
    num_samples=150,
    model="mistralai/mistral-7b-instruct"
)

```

## How the Generation Process Works

Understanding the internal flow helps debug issues and optimize output quality.

### Step 1: Prompt Construction

The `_openrouter` function assembles a payload with:

- A system message instructing the model to *"generate training data for a tool-calling and extraction model"*
- The formatted `_GEN_TEMPLATE` injected with your tool schemas
- Sampling parameters (`temperature`, `max_tokens`)

### Step 2: Response Parsing

Generated content passes through `_parse_array`, which extracts the JSON array from the LLM response and validates basic structure.

### Step 3: Deduplication

The `generate_dataset` function computes a `_dedup_key` from `(query, answers)` tuples to prevent identical examples from entering the final dataset. This runs continuously during generation, so requesting 1,000 samples may trigger slightly more API calls to compensate for duplicates.

### Step 4: Concurrency Management

`ThreadPoolExecutor` manages parallel workers. Each worker independently calls `generate_examples`, and results aggregate into a thread-safe collection with progress reporting.

## Customizing Model Behavior

### Selecting Different Models

OpenRouter provides access to dozens of models. Benchmark your use case:

```python

# Cost-effective option with good tool-calling performance

samples = generate_dataset(tools, num_samples=500, model="mistralai/mistral-7b-instruct")

# Higher quality, higher cost

samples = generate_dataset(tools, num_samples=500, model="anthropic/claude-3-sonnet")

```

### Adjusting Temperature

For synthetic training data, moderate temperatures (0.5-0.8) typically balance diversity and coherence:

```python

# Conservative, predictable outputs

samples = generate_dataset(tools, num_samples=200, temperature=0.3)

# Creative, varied examples (may require more filtering)

samples = generate_dataset(tools, num_samples=200, temperature=1.2)

```

## Error Handling and Edge Cases

The `_openrouter` implementation includes basic request/response handling, but production usage should account for:

- **Rate limiting** – Reduce `--workers` or implement backoff if hitting OpenRouter limits.
- **Malformed JSON** – Enable verbose logging to inspect raw LLM outputs when `_parse_array` fails.
- **Insufficient diversity** – Increase `temperature` or switch models if deduplication removes too many samples.

## Summary

- **Primary entry points**: `needle generate-data` CLI or `generate_dataset()` / `augment_jsonl()` from [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py).
- **Configuration**: Set `OPENROUTER_API_KEY`; optionally override `OPENROUTER_URL`.
- **Key files**: [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py) contains all generation logic; [`needle/cli.py`](https://github.com/cactus-compute/needle/blob/main/needle/cli.py) provides the command-line interface.
- **Concurrency**: Use `--workers` or `workers=` to parallelize API calls for faster generation.
- **Deduplication**: Automatic based on `(query, answers)` keys during dataset assembly.
- **Flexibility**: Any OpenRouter-compatible model and sampling parameters are supported.

## Frequently Asked Questions

### What file format should my tool schemas use?

Your tool definitions should be in standard JSON format—typically an array of objects with `name`, `description`, and `parameters` fields, matching the JSON Schema convention used by OpenAI's function calling. The `generate_dataset` function passes these schemas directly into the prompt template without modification.

### Why is my generation slower than expected?

OpenRouter rate limits and your `--workers` setting are the primary factors. Reduce workers if encountering HTTP 429 errors, or increase them (up to 20-50) for faster throughput on uncapped tiers. Network latency and model response time also vary—smaller models like `mistral-7b-instruct` are generally faster than `claude-3-opus`.

### Can I resume an interrupted generation job?

The current implementation in [`needle/model/finetune.py`](https://github.com/cactus-compute/needle/blob/main/needle/model/finetune.py) does not support checkpointing. Use `augment_jsonl` as a workaround: save intermediate results periodically, then resume by pointing `--augment` at the last successful output file. A proper resume feature would require tracking completed `(query, answers)` pairs in a sidecar file.