How to Configure Reward Functions Using Soup Reward Synth: A Complete Guide

Use soup reward synth to automatically generate deterministic reward verifiers from reference data through a four-stage pipeline: detection, induction, rendering, and calibration.

The soup reward synth command in the Soup repository (MakazhanAlpamys/Soup) provides a zero-configuration way to create verifiable reward functions for LLM training. By analyzing your gold-standard outputs, it detects the appropriate verifier family, synthesizes a concrete specification, renders a self-contained Python module, and calibrates it against positive and negative examples before emitting anything.

The Four-Stage Reward Synthesis Pipeline

The synthesis engine lives in src/soup_cli/utils/reward_synth.py. Each stage is designed to be deterministic and auditable, as noted in the pipeline comment at lines 8-10.

Stage 1: Detect the Verifier Family

The detect_kind function scans your reference strings and assigns one of four supported families based on a confidence threshold defined by _MIN_CONFIDENCE (lines 78-99):

  • numeric – All gold values parse as numbers
  • json_schema – All gold values are valid JSON objects
  • regex – Gold strings share consistent length and character patterns
  • tool_call – Gold values represent function/tool invocations

If detection falls below the confidence threshold, the command aborts with guidance.

Stage 2: Induce a Concrete Specification

Each family has a dedicated inducer that builds a typed specification:

  • induce_numeric (lines 4-20): Creates a NumericSpec with configurable tolerance (_DEFAULT_NUMERIC_TOLERANCE)
  • induce_json_schema (lines 38-74): Derives a top-level JSON schema from your gold objects
  • induce_regex (lines 13-33): Generates a strict regex pattern when all references share length constraints
  • induce_tool_call (lines 76-98): Produces a ToolCallSpec tracking required and allowed argument keys per tool name

These specs are pure Python dataclasses, not opaque models, so you can inspect and modify them.

Stage 3: Render a Self-Contained Verifier

The render_verifier_py function (lines 18-50) combines your spec with a template consisting of:

  • _HEADER – Shared imports and helper utilities
  • Family-specific body – _NUMERIC_BODY, _JSON_SCHEMA_BODY, _REGEX_BODY, or _TOOL_CALL_BODY

The output is a single importable Python file exposing reward_fn(completions, **kwargs). You can extend this file manually or regenerate it as your data evolves.

Stage 4: Calibrate Against Positives and Negatives

Before finalizing, the CLI validates the verifier using perturb_negatives (lines 55-92) to generate corrupted variants of your references. The resulting CalibrationReport enforces three safety checks:

  1. Self-acceptance – Must accept at least _MIN_SELF_ACCEPT fraction of original golds
  2. Discrimination > 0 – Must score positives strictly higher than negatives
  3. User threshold – Must meet or exceed --min-discrimination (default 0.5)

Failure on any check aborts with exit code 2 and no verifier is emitted, preventing degenerate reward functions from entering your training loop.

CLI Usage and Examples

The command implementation resides in src/soup_cli/commands/reward.py (lines 56-98). It handles JSONL parsing, orchestrates the synthesis library, manages secure file writes via atomic_write_text, and presents results through Rich console panels.

Basic Synthesis (Auto-Detect Kind)

soup reward synth references.jsonl -o reward.py

Your references.jsonl should contain one JSON object per line with at minimum an "output" field. The command infers everything else.

Force Numeric Verifier with Custom Tolerance

soup reward synth references.jsonl -o reward.py --kind numeric --tolerance 1e-4

Use --kind to override auto-detection when you know your data type. Tolerance controls how close predicted numbers must be to gold values.

Preview Without Writing

soup reward synth references.jsonl --plan-only

Outputs a _spec_summary showing the detected family, induced parameters, and expected discrimination—useful for debugging data quality issues.

Stricter Discrimination Threshold

soup reward synth references.jsonl -o reward.py --min-discrimination 0.7

Raises the bar for positive-negative separation. Higher values produce more conservative verifiers that may reject borderline correct outputs.

Overwrite Existing Verifier

soup reward synth references.jsonl -o reward.py --force

By default, the CLI refuses to clobber existing files. Use --force when iterating on your reference set.

Generated Verifier Structure

The output file is designed for readability and editability. Here are excerpts from two common families:

Numeric Verifier

def reward_fn(completions, **kwargs):
    answers = kwargs.get("answer", [])
    out = []
    for completion, expected in zip(completions, answers):
        predicted = _extract_number(_last_content(completion))
        gold = str(expected).strip()
        out.append(1.0 if _numbers_match(predicted, gold, _TOLERANCE) else 0.0)
    return out

The _TOLERANCE constant is set from your --tolerance flag or the default. Helper functions _extract_number and _numbers_match handle formatting variations.

JSON Schema Verifier

def reward_fn(completions, **kwargs):
    out = []
    for completion in completions:
        content = _last_content(completion)
        fenced = re.search(r"```(?:json)?\s*(.*?)```", content, re.DOTALL)
        if fenced:
            content = fenced.group(1)
        try:
            data = json.loads(content.strip())
        except (json.JSONDecodeError, ValueError):
            out.append(0.0)
            continue
        out.append(1.0 if _matches_schema(data) else 0.0)
    return out

This handles both raw JSON and fenced code blocks. The _matches_schema function validates against the induced schema from your gold objects.

Loading and Using Your Verifier

Once generated, import your reward function through the Soup trainer utilities:

from soup_cli.trainer.rewards import load_reward_fn

reward_fn = load_reward_fn("reward.py")
scores = reward_fn(completions, answer=reference_answers)

The load_reward_fn utility (in src/soup_cli/trainer/rewards.py) executes the module safely and validates the required signature before returning the callable.

Security and File Handling

The CLI uses helpers from src/soup_cli/utils/paths.py to enforce:

  • enforce_under_cwd_and_no_symlink – Prevents path traversal and symlink attacks on output paths
  • atomic_write_text – Ensures verifiers are written completely or not at all, avoiding partial files

These guards matter when running synthesis in automated CI pipelines or shared environments.

Summary

  • soup reward synth converts reference data into executable, deterministic reward functions through detection, induction, rendering, and calibration
  • Four verifier families cover numeric, JSON schema, regex, and tool call output types
  • The calibration stage prevents emission of reward functions that fail to discriminate positives from negatives
  • Generated files are pure Python, editable, and loadable via load_reward_fn
  • Security helpers ensure safe file operations even in untrusted environments

Frequently Asked Questions

What input format does soup reward synth require?

The command accepts a JSONL file where each line is a JSON object containing at minimum an "output" field with your gold-standard reference. Additional metadata fields are preserved but not required. The CLI parser _read_jsonl in src/soup_cli/commands/reward.py validates structure before passing to the synthesis engine.

Can I modify a generated verifier after synthesis?

Yes. The rendered Python file is self-contained and intended for human editing. You can adjust tolerances, relax schema constraints, or add custom preprocessing. Regenerate with --force if your reference data changes significantly, or maintain the file manually for fine-grained control.

Why did my synthesis fail with exit code 2?

Exit code 2 indicates calibration failure. Your verifier either rejected too many gold references, failed to score positives above negatives, or fell below your --min-discrimination threshold. This protects your training pipeline from reward hacking. Inspect with --plan-only to see the induced spec, then add more diverse references or adjust tolerance/schema parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →