How to Implement Data Quality Scoring with `soup data score`

The soup data score command runs a composite scorecard that evaluates JSONL datasets for PII, toxicity, language distribution, educational value, and benchmark contamination—all without mandatory dependencies.

MakazhanAlpamys/Soup provides a zero-dependency pipeline for data quality assessment. Whether you're preparing training data for language models or auditing existing datasets, the soup data score toolkit offers both a convenient CLI interface and programmatic APIs for deep inspection. This guide walks through all implementation methods based on the actual source code in the repository.

Core Scoring Metrics Explained

The scoring system aggregates six heuristic categories defined in src/soup_cli/utils/data_score.py. Each metric serves a distinct quality assurance purpose.

PII Detection

The detect_pii function scans text for personally identifiable information using fast regex patterns:

  • Email addresses
  • Phone numbers
  • Social Security Numbers
  • Credit card numbers

When the optional [data-pro] extra is installed, the system can leverage Microsoft Presidio for enhanced detection. Without it, the regex baseline still catches common patterns efficiently.

Toxicity Scoring

The score_toxicity function implements a keyword-based classifier that flags hateful or violent language. While lightweight, this heuristic provides rapid triage for problematic content.

Language Identification

detect_language attempts ISO-639-1 code assignment through two strategies:

  1. Primary: langdetect package (optional dependency)
  2. Fallback: stop-word frequency heuristic

This dual approach ensures functionality even in minimal installations.

Educational Value Assessment

score_educational_value quantifies pedagogical utility by analyzing:

  • Text length (sufficient content depth)
  • Lexical diversity (vocabulary richness)

Higher scores indicate content more suitable for training or instructional use.

Decontamination

decontaminate_rows removes training data that overlaps with public benchmarks (MMLU, GSM8K, etc.). It uses n-gram overlap logic to prevent data leakage into evaluation sets.

Aggregate Reporting

compute_scorecard orchestrates all metrics into a unified ScoreReport dataclass with totals, flagged counts, and distributions.

CLI Implementation

The command-line interface in src/soup_cli/commands/data_score.py provides the fastest path to data quality assessment.

Basic Scoring Command

soup data score \
    --input training.jsonl \
    --benchmarks mmlu,gsm8k \
    --threshold 0.85

Parameters:

  • --input: Path to JSONL file (must be under current working directory for security)
  • --benchmarks: Comma-separated list of benchmark corpora for decontamination
  • --threshold: N-gram overlap threshold for contamination detection (0.0–1.0)

The output renders as a Rich table showing:

  • Total row count
  • PII-flagged rows
  • Toxic-flagged rows
  • Mean educational score
  • Decontaminated row count
  • Language distribution histogram

Security Constraints

The CLI enforces path containment via is_under_cwd (defined in src/soup_cli/utils/paths.py) to prevent accidental filesystem escape when processing datasets.

Programmatic Implementation

For integration into larger pipelines, import functions directly from src/soup_cli/utils/data_score.py.

Full Scorecard Computation

from soup_cli.utils.data_score import (
    load_jsonl_rows,
    compute_scorecard,
    BENCHMARKS,
)

# Load dataset with validation and size limits

rows = load_jsonl_rows("training.jsonl")

# Select benchmarks for decontamination

benchmarks = ["mmlu", "gsm8k"]  # keys must exist in BENCHMARKS dict

# Generate composite report

report = compute_scorecard(
    rows,
    benchmarks=benchmarks,
    decontaminate_threshold=0.9,
)

print(f"Total rows: {report.total}")
print(f"PII flagged: {report.pii_flagged}")
print(f"Toxic flagged: {report.toxic_flagged}")
print(f"Mean educational score: {report.educational_mean:.3f}")
print(f"Languages detected: {report.language_counts}")

Manual Decontamination

When you need fine-grained control over the removal process:

from soup_cli.utils.data_score import (
    load_jsonl_rows,
    decontaminate_rows,
    extract_row_text,
)

rows = load_jsonl_rows("raw.jsonl")
benchmark_texts = [
    "A bat and a ball cost together",
    "What is the capital of France?"
]

kept, removed = decontaminate_rows(
    rows,
    benchmark_texts,
    n=8,           # n-gram size for comparison

    threshold=0.8,  # overlap threshold

)

print(f"Kept {len(kept)} rows, removed {len(removed)}")

# Process kept rows downstream

Single-Record PII Scanning

For real-time or streaming applications:

from soup_cli.utils.data_score import detect_pii

hits = detect_pii("Contact me at alice@example.com or +1-555-123-4567.")
print(hits)

# [{'kind': 'email', 'snippet': 'alice@example.com'},

#  {'kind': 'phone', 'snippet': '+1-555-123-4567'}]

Lazy Import Architecture

The codebase respects lazy-import principles. Heavy optional dependencies—Presidio for advanced PII detection, langdetect for language identification—are imported only when their specific functions are invoked. This keeps startup time minimal and prevents installation failures in constrained environments.

Supported Benchmarks for Decontamination

The BENCHMARKS dictionary in src/soup_cli/utils/data_score.py maps standard evaluation corpora to their n-gram signatures. Current implementations include:

  • mmlu: Massive Multitask Language Understanding
  • gsm8k: Grade School Math 8K
  • hellaswag: Commonsense reasoning evaluation
  • arc: AI2 Reasoning Challenge

Extend this dictionary for custom benchmark integration.

Summary

  • soup data score combines PII detection, toxicity scoring, language identification, educational assessment, and benchmark decontamination in one command
  • Zero-dependency core keeps the toolkit lightweight; optional extras provide enhanced capabilities
  • CLI entry point: src/soup_cli/commands/data_score.py with Rich table output
  • Core logic: src/soup_cli/utils/data_score.py containing detect_pii, score_toxicity, detect_language, score_educational_value, decontaminate_rows, and compute_scorecard
  • Security: Path containment via is_under_cwd prevents directory traversal
  • Programmatic API: All functions importable for custom pipeline integration

Frequently Asked Questions

What file formats does soup data score accept?

The toolkit exclusively processes JSON Lines (JSONL) files where each line is a valid JSON object. The load_jsonl_rows helper validates structure and applies configurable size limits to prevent memory exhaustion with large datasets.

Can I use the scoring functions without installing optional dependencies?

Yes. The core functionality operates without Presidio, langdetect, or other extras. Fallback implementations—regex-based PII detection and stop-word language heuristics—maintain full functionality with reduced precision compared to their ML-enhanced counterparts.

How do I add custom benchmarks for decontamination?

The BENCHMARKS dictionary in src/soup_cli/utils/data_score.py accepts new entries mapping benchmark names to n-gram corpus sets. After adding your benchmark texts, reference them by key in the benchmarks parameter of compute_scorecard or via --benchmarks in the CLI.

What threshold should I use for decontamination?

The default threshold of 0.85 balances aggressive contamination removal against false positives. Higher thresholds (0.9–0.95) preserve more data but risk benchmark leakage; lower thresholds (0.7–0.8) maximize cleanliness at the cost of dataset size. Tune based on your downstream evaluation sensitivity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →