How to Implement Data Quality Scoring with `soup data score`
The soup data score command runs a composite scorecard that evaluates JSONL datasets for PII, toxicity, language distribution, educational value, and benchmark contamination—all without mandatory dependencies.
MakazhanAlpamys/Soup provides a zero-dependency pipeline for data quality assessment. Whether you're preparing training data for language models or auditing existing datasets, the soup data score toolkit offers both a convenient CLI interface and programmatic APIs for deep inspection. This guide walks through all implementation methods based on the actual source code in the repository.
Core Scoring Metrics Explained
The scoring system aggregates six heuristic categories defined in src/soup_cli/utils/data_score.py. Each metric serves a distinct quality assurance purpose.
PII Detection
The detect_pii function scans text for personally identifiable information using fast regex patterns:
- Email addresses
- Phone numbers
- Social Security Numbers
- Credit card numbers
When the optional [data-pro] extra is installed, the system can leverage Microsoft Presidio for enhanced detection. Without it, the regex baseline still catches common patterns efficiently.
Toxicity Scoring
The score_toxicity function implements a keyword-based classifier that flags hateful or violent language. While lightweight, this heuristic provides rapid triage for problematic content.
Language Identification
detect_language attempts ISO-639-1 code assignment through two strategies:
- Primary:
langdetectpackage (optional dependency) - Fallback: stop-word frequency heuristic
This dual approach ensures functionality even in minimal installations.
Educational Value Assessment
score_educational_value quantifies pedagogical utility by analyzing:
- Text length (sufficient content depth)
- Lexical diversity (vocabulary richness)
Higher scores indicate content more suitable for training or instructional use.
Decontamination
decontaminate_rows removes training data that overlaps with public benchmarks (MMLU, GSM8K, etc.). It uses n-gram overlap logic to prevent data leakage into evaluation sets.
Aggregate Reporting
compute_scorecard orchestrates all metrics into a unified ScoreReport dataclass with totals, flagged counts, and distributions.
CLI Implementation
The command-line interface in src/soup_cli/commands/data_score.py provides the fastest path to data quality assessment.
Basic Scoring Command
soup data score \
--input training.jsonl \
--benchmarks mmlu,gsm8k \
--threshold 0.85
Parameters:
--input: Path to JSONL file (must be under current working directory for security)--benchmarks: Comma-separated list of benchmark corpora for decontamination--threshold: N-gram overlap threshold for contamination detection (0.0–1.0)
The output renders as a Rich table showing:
- Total row count
- PII-flagged rows
- Toxic-flagged rows
- Mean educational score
- Decontaminated row count
- Language distribution histogram
Security Constraints
The CLI enforces path containment via is_under_cwd (defined in src/soup_cli/utils/paths.py) to prevent accidental filesystem escape when processing datasets.
Programmatic Implementation
For integration into larger pipelines, import functions directly from src/soup_cli/utils/data_score.py.
Full Scorecard Computation
from soup_cli.utils.data_score import (
load_jsonl_rows,
compute_scorecard,
BENCHMARKS,
)
# Load dataset with validation and size limits
rows = load_jsonl_rows("training.jsonl")
# Select benchmarks for decontamination
benchmarks = ["mmlu", "gsm8k"] # keys must exist in BENCHMARKS dict
# Generate composite report
report = compute_scorecard(
rows,
benchmarks=benchmarks,
decontaminate_threshold=0.9,
)
print(f"Total rows: {report.total}")
print(f"PII flagged: {report.pii_flagged}")
print(f"Toxic flagged: {report.toxic_flagged}")
print(f"Mean educational score: {report.educational_mean:.3f}")
print(f"Languages detected: {report.language_counts}")
Manual Decontamination
When you need fine-grained control over the removal process:
from soup_cli.utils.data_score import (
load_jsonl_rows,
decontaminate_rows,
extract_row_text,
)
rows = load_jsonl_rows("raw.jsonl")
benchmark_texts = [
"A bat and a ball cost together",
"What is the capital of France?"
]
kept, removed = decontaminate_rows(
rows,
benchmark_texts,
n=8, # n-gram size for comparison
threshold=0.8, # overlap threshold
)
print(f"Kept {len(kept)} rows, removed {len(removed)}")
# Process kept rows downstream
Single-Record PII Scanning
For real-time or streaming applications:
from soup_cli.utils.data_score import detect_pii
hits = detect_pii("Contact me at alice@example.com or +1-555-123-4567.")
print(hits)
# [{'kind': 'email', 'snippet': 'alice@example.com'},
# {'kind': 'phone', 'snippet': '+1-555-123-4567'}]
Lazy Import Architecture
The codebase respects lazy-import principles. Heavy optional dependencies—Presidio for advanced PII detection, langdetect for language identification—are imported only when their specific functions are invoked. This keeps startup time minimal and prevents installation failures in constrained environments.
Supported Benchmarks for Decontamination
The BENCHMARKS dictionary in src/soup_cli/utils/data_score.py maps standard evaluation corpora to their n-gram signatures. Current implementations include:
- mmlu: Massive Multitask Language Understanding
- gsm8k: Grade School Math 8K
- hellaswag: Commonsense reasoning evaluation
- arc: AI2 Reasoning Challenge
Extend this dictionary for custom benchmark integration.
Summary
soup data scorecombines PII detection, toxicity scoring, language identification, educational assessment, and benchmark decontamination in one command- Zero-dependency core keeps the toolkit lightweight; optional extras provide enhanced capabilities
- CLI entry point:
src/soup_cli/commands/data_score.pywith Rich table output - Core logic:
src/soup_cli/utils/data_score.pycontainingdetect_pii,score_toxicity,detect_language,score_educational_value,decontaminate_rows, andcompute_scorecard - Security: Path containment via
is_under_cwdprevents directory traversal - Programmatic API: All functions importable for custom pipeline integration
Frequently Asked Questions
What file formats does soup data score accept?
The toolkit exclusively processes JSON Lines (JSONL) files where each line is a valid JSON object. The load_jsonl_rows helper validates structure and applies configurable size limits to prevent memory exhaustion with large datasets.
Can I use the scoring functions without installing optional dependencies?
Yes. The core functionality operates without Presidio, langdetect, or other extras. Fallback implementations—regex-based PII detection and stop-word language heuristics—maintain full functionality with reduced precision compared to their ML-enhanced counterparts.
How do I add custom benchmarks for decontamination?
The BENCHMARKS dictionary in src/soup_cli/utils/data_score.py accepts new entries mapping benchmark names to n-gram corpus sets. After adding your benchmark texts, reference them by key in the benchmarks parameter of compute_scorecard or via --benchmarks in the CLI.
What threshold should I use for decontamination?
The default threshold of 0.85 balances aggressive contamination removal against false positives. Higher thresholds (0.9–0.95) preserve more data but risk benchmark leakage; lower thresholds (0.7–0.8) maximize cleanliness at the cost of dataset size. Tune based on your downstream evaluation sensitivity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →