# How to Implement Data Quality Scoring with `soup data score`

> Implement data quality scoring with soup data score. Evaluate PII toxicity language distribution educational value and benchmark contamination in JSONL datasets without dependencies.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-08-16

---

**The `soup data score` command runs a composite scorecard that evaluates JSONL datasets for PII, toxicity, language distribution, educational value, and benchmark contamination—all without mandatory dependencies.**

MakazhanAlpamys/Soup provides a zero-dependency pipeline for data quality assessment. Whether you're preparing training data for language models or auditing existing datasets, the `soup data score` toolkit offers both a convenient CLI interface and programmatic APIs for deep inspection. This guide walks through all implementation methods based on the actual source code in the repository.

## Core Scoring Metrics Explained

The scoring system aggregates six heuristic categories defined in [`src/soup_cli/utils/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_score.py). Each metric serves a distinct quality assurance purpose.

### PII Detection

The `detect_pii` function scans text for personally identifiable information using fast regex patterns:

- **Email addresses**
- **Phone numbers**
- **Social Security Numbers**
- **Credit card numbers**

When the optional `[data-pro]` extra is installed, the system can leverage Microsoft Presidio for enhanced detection. Without it, the regex baseline still catches common patterns efficiently.

### Toxicity Scoring

The `score_toxicity` function implements a keyword-based classifier that flags hateful or violent language. While lightweight, this heuristic provides rapid triage for problematic content.

### Language Identification

`detect_language` attempts ISO-639-1 code assignment through two strategies:

1. Primary: `langdetect` package (optional dependency)
2. Fallback: stop-word frequency heuristic

This dual approach ensures functionality even in minimal installations.

### Educational Value Assessment

`score_educational_value` quantifies pedagogical utility by analyzing:

- **Text length** (sufficient content depth)
- **Lexical diversity** (vocabulary richness)

Higher scores indicate content more suitable for training or instructional use.

### Decontamination

`decontaminate_rows` removes training data that overlaps with public benchmarks (MMLU, GSM8K, etc.). It uses n-gram overlap logic to prevent data leakage into evaluation sets.

### Aggregate Reporting

`compute_scorecard` orchestrates all metrics into a unified `ScoreReport` dataclass with totals, flagged counts, and distributions.

## CLI Implementation

The command-line interface in [`src/soup_cli/commands/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/data_score.py) provides the fastest path to data quality assessment.

### Basic Scoring Command

```bash
soup data score \
    --input training.jsonl \
    --benchmarks mmlu,gsm8k \
    --threshold 0.85

```

**Parameters:**

- `--input`: Path to JSONL file (must be under current working directory for security)
- `--benchmarks`: Comma-separated list of benchmark corpora for decontamination
- `--threshold`: N-gram overlap threshold for contamination detection (0.0–1.0)

The output renders as a Rich table showing:

- Total row count
- PII-flagged rows
- Toxic-flagged rows
- Mean educational score
- Decontaminated row count
- Language distribution histogram

### Security Constraints

The CLI enforces **path containment** via `is_under_cwd` (defined in [`src/soup_cli/utils/paths.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/paths.py)) to prevent accidental filesystem escape when processing datasets.

## Programmatic Implementation

For integration into larger pipelines, import functions directly from [`src/soup_cli/utils/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_score.py).

### Full Scorecard Computation

```python
from soup_cli.utils.data_score import (
    load_jsonl_rows,
    compute_scorecard,
    BENCHMARKS,
)

# Load dataset with validation and size limits

rows = load_jsonl_rows("training.jsonl")

# Select benchmarks for decontamination

benchmarks = ["mmlu", "gsm8k"]  # keys must exist in BENCHMARKS dict

# Generate composite report

report = compute_scorecard(
    rows,
    benchmarks=benchmarks,
    decontaminate_threshold=0.9,
)

print(f"Total rows: {report.total}")
print(f"PII flagged: {report.pii_flagged}")
print(f"Toxic flagged: {report.toxic_flagged}")
print(f"Mean educational score: {report.educational_mean:.3f}")
print(f"Languages detected: {report.language_counts}")

```

### Manual Decontamination

When you need fine-grained control over the removal process:

```python
from soup_cli.utils.data_score import (
    load_jsonl_rows,
    decontaminate_rows,
    extract_row_text,
)

rows = load_jsonl_rows("raw.jsonl")
benchmark_texts = [
    "A bat and a ball cost together",
    "What is the capital of France?"
]

kept, removed = decontaminate_rows(
    rows,
    benchmark_texts,
    n=8,           # n-gram size for comparison

    threshold=0.8,  # overlap threshold

)

print(f"Kept {len(kept)} rows, removed {len(removed)}")

# Process kept rows downstream

```

### Single-Record PII Scanning

For real-time or streaming applications:

```python
from soup_cli.utils.data_score import detect_pii

hits = detect_pii("Contact me at alice@example.com or +1-555-123-4567.")
print(hits)

# [{'kind': 'email', 'snippet': 'alice@example.com'},

#  {'kind': 'phone', 'snippet': '+1-555-123-4567'}]

```

## Lazy Import Architecture

The codebase respects **lazy-import** principles. Heavy optional dependencies—Presidio for advanced PII detection, `langdetect` for language identification—are imported only when their specific functions are invoked. This keeps startup time minimal and prevents installation failures in constrained environments.

## Supported Benchmarks for Decontamination

The `BENCHMARKS` dictionary in [`src/soup_cli/utils/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_score.py) maps standard evaluation corpora to their n-gram signatures. Current implementations include:

- **mmlu**: Massive Multitask Language Understanding
- **gsm8k**: Grade School Math 8K
- **hellaswag**: Commonsense reasoning evaluation
- **arc**: AI2 Reasoning Challenge

Extend this dictionary for custom benchmark integration.

## Summary

- **`soup data score`** combines PII detection, toxicity scoring, language identification, educational assessment, and benchmark decontamination in one command
- **Zero-dependency core** keeps the toolkit lightweight; optional extras provide enhanced capabilities
- **CLI entry point**: [`src/soup_cli/commands/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/data_score.py) with Rich table output
- **Core logic**: [`src/soup_cli/utils/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_score.py) containing `detect_pii`, `score_toxicity`, `detect_language`, `score_educational_value`, `decontaminate_rows`, and `compute_scorecard`
- **Security**: Path containment via `is_under_cwd` prevents directory traversal
- **Programmatic API**: All functions importable for custom pipeline integration

## Frequently Asked Questions

### What file formats does `soup data score` accept?

The toolkit exclusively processes **JSON Lines (JSONL)** files where each line is a valid JSON object. The `load_jsonl_rows` helper validates structure and applies configurable size limits to prevent memory exhaustion with large datasets.

### Can I use the scoring functions without installing optional dependencies?

Yes. The core functionality operates without Presidio, `langdetect`, or other extras. Fallback implementations—regex-based PII detection and stop-word language heuristics—maintain full functionality with reduced precision compared to their ML-enhanced counterparts.

### How do I add custom benchmarks for decontamination?

The `BENCHMARKS` dictionary in [`src/soup_cli/utils/data_score.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_score.py) accepts new entries mapping benchmark names to n-gram corpus sets. After adding your benchmark texts, reference them by key in the `benchmarks` parameter of `compute_scorecard` or via `--benchmarks` in the CLI.

### What threshold should I use for decontamination?

The default threshold of **0.85** balances aggressive contamination removal against false positives. Higher thresholds (0.9–0.95) preserve more data but risk benchmark leakage; lower thresholds (0.7–0.8) maximize cleanliness at the cost of dataset size. Tune based on your downstream evaluation sensitivity.