How to Create Synthetic Training Data with Soup Data Forge: A Complete Guide

Soup data forge is a multi‑stage pipeline that converts raw documents into high‑quality synthetic training rows while preserving full provenance through chunking, LLM judging, and uncertainty‑based pruning.

The soup data forge command is the primary entry point for generating supervised fine‑tuning (SFT) datasets from unstructured text. As implemented in MakazhanAlpamys/Soup, it combines document chunking, active‑learning filtering, and provenance tracking into a single reproducible workflow.

Overview of the Soup Data Forge Pipeline

The pipeline consists of three tightly integrated stages defined in src/soup_cli/utils/data_forge.py:

  • Document discovery and chunking
  • Judging with uncertainty‑based pruning
  • Provenance manifest generation

Each stage enforces strict security boundaries: path containment via os.lstat with S_ISLNK checks, atomic writes using tempfile → os.replace, and SSRF‑hardening for remote judge providers.

Stage 1: Document Chunking

The forge command discovers all supported text files (.txt, .md, .json, .jsonl) one level deep under the --docs directory. Each file is split into chunks using the same tokenizer that will later be used for training.

Key implementation details from src/soup_cli/utils/data_forge.py:

  • Default chunk size: approximately 256 tokens
  • CWD‑contained execution: the process refuses to follow symlinks or traverse dot‑directories
  • Path‑traversal protection prevents reading files outside the designated document root

# Prepare a directory with source documents

mkdir -p my_docs
cp *.md *.txt my_docs/

# Run chunking as part of the full pipeline

soup data forge \
  --docs ./my_docs/ \
  --task sft \
  --target-rows 1000 \
  --output forge_dataset.jsonl

Stage 2: Judging and Active‑Learning Pruning

For every chunk, the pipeline invokes an LLM judge specified by --judge‑provider (ollama, anthropic, or vllm). The judge generates a synthetic response that is compared to the source chunk using a Jaccard‑style similarity metric.

Rows with similarity above --uncertainty‑threshold are pruned as "too easy." The remaining uncertain rows—those where the judge's output diverges meaningfully from the source—are retained because they provide the strongest learning signal.

This active‑learning approach reduces dataset noise and improves downstream model performance. The threshold is configurable; lower values retain more diverse, challenging examples.


# Retain only rows with judge‑source uncertainty below 0.4

soup data forge \
  --docs ./my_docs/ \
  --task sft \
  --target-rows 1000 \
  --uncertainty-threshold 0.4 \
  --judge-provider ollama \
  --judge-model llama3 \
  --output forge_dataset.jsonl \
  --provenance forge_provenance.json

Stage 3: Provenance Manifest Generation

Every retained synthetic row is written to the output JSONL (--output). Simultaneously, a separate provenance file (--provenance) creates a complete audit trail mapping:

  • Row ID → originating document path
  • Judge call ID and model version
  • Chunk ID within the source file
  • Filter score (uncertainty metric)

This manifest enables compliance verification, debugging, and full reproducibility of training datasets. According to the Soup source code, provenance files follow a strict schema documented in docs/data.md.

Complete End‑to‑End Workflow

The following example mirrors the official Synthetic Data Workflow from examples/synthetic_workflow.md:


# 1. Generate initial synthetic data with a prompt

soup data generate \
  --prompt "Write concise Python instruction/response pairs." \
  --provider ollama \
  --model llama3.2:3b \
  --topic "Python error handling" \
  --count 200 \
  --output ./synth_raw.jsonl

# 2. Filter for coherence

soup data filter ./synth_raw.jsonl \
  --output ./synth_filtered.jsonl \
  --min-coherence 0.5

# 3. Score data quality

soup data score --input ./synth_filtered.jsonl

# 4. Remove benchmark contamination

soup data decontaminate \
  --input ./synth_filtered.jsonl \
  --output ./synth_clean.jsonl \
  --benchmarks mmlu,gsm8k

# 5. Train with the bundled recipe

soup train --config examples/synthetic_workflow.yaml --yes

Key Implementation Files

Path Purpose
src/soup_cli/commands/data_forge.py CLI entry point defining soup data forge arguments and validation
src/soup_cli/utils/data_forge.py Core logic: chunking, judging, uncertainty pruning, provenance writing
examples/synthetic_workflow.yaml Training configuration referencing synthetic datasets
examples/synthetic_workflow.md Narrative documentation for the full pipeline
docs/data.md Flag reference and provenance schema specification

Task Types and Customization

The --task parameter controls output format:

  • sft (default): instruction‑response pairs for supervised fine‑tuning
  • preference: prompt with chosen/rejected responses for RLHF
  • tool: function‑calling examples with tool definitions

All tasks share the same chunking and judging infrastructure in src/soup_cli/utils/data_forge.py.

Inspecting and Validating Results

After generation, use built‑in utilities to verify dataset quality:


# Quick statistics: row count, token distribution

soup data inspect forge_dataset.jsonl

# Run full quality scorecard

soup data score --input forge_dataset.jsonl

Summary

  • Soup data forge transforms raw documents into training data through three stages: chunking, judging with uncertainty pruning, and provenance tracking.
  • The active‑learning filter removes rows that are too easy for the judge, preserving only high‑signal examples.
  • Provenance manifests enable full audit trails and regulatory compliance.
  • Security features include path containment, atomic writes, and SSRF‑hardening throughout src/soup_cli/utils/data_forge.py.

Frequently Asked Questions

What file formats does soup data forge support?

soup data forge processes .txt, .md, .json, and .jsonl files. Discovery is limited to one level deep under the --docs directory, with automatic exclusion of dot‑files and symlinked paths to prevent traversal attacks.

How does the uncertainty threshold affect my dataset?

The --uncertainty-threshold (default varies by provider) sets the Jaccard similarity cutoff. Rows where judge output and source text are more similar than this threshold are discarded. Lower thresholds (e.g., 0.3) keep harder, more diverse examples; higher thresholds (e.g., 0.6) retain easier, more faithful reconstructions.

Can I use cloud LLMs as judges?

Yes. The --judge-provider flag accepts anthropic for Claude models, vllm for self‑hosted endpoints, or ollama for local inference. All remote providers implement SSRF‑hardening as enforced in src/soup_cli/utils/data_forge.py.

Is the provenance file required?

No, but strongly recommended. Omitting --provenance disables audit trailing. The provenance JSON maps every synthetic row to its source document, chunk ID, and filter score—essential for debugging, compliance, and reproducible ML pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →