# How to Create Synthetic Training Data with Soup Data Forge: A Complete Guide

> Learn to create synthetic training data with Soup Data Forge. This guide details its multi-stage pipeline for high-quality data generation, preserving provenance.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**Soup data forge is a multi‑stage pipeline that converts raw documents into high‑quality synthetic training rows while preserving full provenance through chunking, LLM judging, and uncertainty‑based pruning.**

The `soup data forge` command is the primary entry point for generating supervised fine‑tuning (SFT) datasets from unstructured text. As implemented in `MakazhanAlpamys/Soup`, it combines document chunking, active‑learning filtering, and provenance tracking into a single reproducible workflow.

## Overview of the Soup Data Forge Pipeline

The pipeline consists of three tightly integrated stages defined in [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py):

- **Document discovery and chunking**
- **Judging with uncertainty‑based pruning**
- **Provenance manifest generation**

Each stage enforces strict security boundaries: path containment via `os.lstat` with `S_ISLNK` checks, atomic writes using `tempfile` → `os.replace`, and SSRF‑hardening for remote judge providers.

## Stage 1: Document Chunking

The forge command discovers all supported text files (`.txt`, `.md`, `.json`, `.jsonl`) one level deep under the `--docs` directory. Each file is split into chunks using the same tokenizer that will later be used for training.

Key implementation details from [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py):

- Default chunk size: approximately **256 tokens**
- **CWD‑contained execution**: the process refuses to follow symlinks or traverse dot‑directories
- Path‑traversal protection prevents reading files outside the designated document root

```bash

# Prepare a directory with source documents

mkdir -p my_docs
cp *.md *.txt my_docs/

# Run chunking as part of the full pipeline

soup data forge \
  --docs ./my_docs/ \
  --task sft \
  --target-rows 1000 \
  --output forge_dataset.jsonl

```

## Stage 2: Judging and Active‑Learning Pruning

For every chunk, the pipeline invokes an LLM judge specified by `--judge‑provider` (`ollama`, `anthropic`, or `vllm`). The judge generates a synthetic response that is compared to the source chunk using a **Jaccard‑style similarity metric**.

Rows with similarity above `--uncertainty‑threshold` are **pruned** as "too easy." The remaining uncertain rows—those where the judge's output diverges meaningfully from the source—are retained because they provide the strongest learning signal.

This active‑learning approach reduces dataset noise and improves downstream model performance. The threshold is configurable; lower values retain more diverse, challenging examples.

```bash

# Retain only rows with judge‑source uncertainty below 0.4

soup data forge \
  --docs ./my_docs/ \
  --task sft \
  --target-rows 1000 \
  --uncertainty-threshold 0.4 \
  --judge-provider ollama \
  --judge-model llama3 \
  --output forge_dataset.jsonl \
  --provenance forge_provenance.json

```

## Stage 3: Provenance Manifest Generation

Every retained synthetic row is written to the output JSONL (`--output`). Simultaneously, a separate provenance file (`--provenance`) creates a complete audit trail mapping:

- Row ID → originating document path
- Judge call ID and model version
- Chunk ID within the source file
- Filter score (uncertainty metric)

This manifest enables compliance verification, debugging, and full reproducibility of training datasets. According to the Soup source code, provenance files follow a strict schema documented in [`docs/data.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/data.md).

## Complete End‑to‑End Workflow

The following example mirrors the official **Synthetic Data Workflow** from [`examples/synthetic_workflow.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/examples/synthetic_workflow.md):

```bash

# 1. Generate initial synthetic data with a prompt

soup data generate \
  --prompt "Write concise Python instruction/response pairs." \
  --provider ollama \
  --model llama3.2:3b \
  --topic "Python error handling" \
  --count 200 \
  --output ./synth_raw.jsonl

# 2. Filter for coherence

soup data filter ./synth_raw.jsonl \
  --output ./synth_filtered.jsonl \
  --min-coherence 0.5

# 3. Score data quality

soup data score --input ./synth_filtered.jsonl

# 4. Remove benchmark contamination

soup data decontaminate \
  --input ./synth_filtered.jsonl \
  --output ./synth_clean.jsonl \
  --benchmarks mmlu,gsm8k

# 5. Train with the bundled recipe

soup train --config examples/synthetic_workflow.yaml --yes

```

## Key Implementation Files

| Path | Purpose |
|------|---------|
| [`src/soup_cli/commands/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/data_forge.py) | CLI entry point defining `soup data forge` arguments and validation |
| [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py) | Core logic: chunking, judging, uncertainty pruning, provenance writing |
| [`examples/synthetic_workflow.yaml`](https://github.com/MakazhanAlpamys/Soup/blob/main/examples/synthetic_workflow.yaml) | Training configuration referencing synthetic datasets |
| [`examples/synthetic_workflow.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/examples/synthetic_workflow.md) | Narrative documentation for the full pipeline |
| [`docs/data.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/data.md) | Flag reference and provenance schema specification |

## Task Types and Customization

The `--task` parameter controls output format:

- **`sft`** (default): instruction‑response pairs for supervised fine‑tuning
- **`preference`**: prompt with chosen/rejected responses for RLHF
- **`tool`**: function‑calling examples with tool definitions

All tasks share the same chunking and judging infrastructure in [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py).

## Inspecting and Validating Results

After generation, use built‑in utilities to verify dataset quality:

```bash

# Quick statistics: row count, token distribution

soup data inspect forge_dataset.jsonl

# Run full quality scorecard

soup data score --input forge_dataset.jsonl

```

## Summary

- **Soup data forge** transforms raw documents into training data through three stages: chunking, judging with uncertainty pruning, and provenance tracking.
- The **active‑learning filter** removes rows that are too easy for the judge, preserving only high‑signal examples.
- **Provenance manifests** enable full audit trails and regulatory compliance.
- Security features include path containment, atomic writes, and SSRF‑hardening throughout [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py).

## Frequently Asked Questions

### What file formats does soup data forge support?

`soup data forge` processes `.txt`, `.md`, `.json`, and `.jsonl` files. Discovery is limited to one level deep under the `--docs` directory, with automatic exclusion of dot‑files and symlinked paths to prevent traversal attacks.

### How does the uncertainty threshold affect my dataset?

The `--uncertainty-threshold` (default varies by provider) sets the Jaccard similarity cutoff. Rows where judge output and source text are **more similar** than this threshold are discarded. Lower thresholds (e.g., 0.3) keep harder, more diverse examples; higher thresholds (e.g., 0.6) retain easier, more faithful reconstructions.

### Can I use cloud LLMs as judges?

Yes. The `--judge-provider` flag accepts `anthropic` for Claude models, `vllm` for self‑hosted endpoints, or `ollama` for local inference. All remote providers implement SSRF‑hardening as enforced in [`src/soup_cli/utils/data_forge.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_forge.py).

### Is the provenance file required?

No, but strongly recommended. Omitting `--provenance` disables audit trailing. The provenance JSON maps every synthetic row to its source document, chunk ID, and filter score—essential for debugging, compliance, and reproducible ML pipelines.