# How to Use Benford's Law Detection in financial_rigor.py for Anomaly Detection

> Detect financial anomalies with Benford's Law detection in financial_rigor.py. Learn how this Python function uses MAD and Chi-square stats to analyze first-digit distributions. Explore the xbtlin/ai-berkshire repo.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The `benford_check` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) analyzes the first-digit distribution of financial datasets to detect anomalies using Mean Absolute Deviation (MAD) and Chi-square statistics against Benford's expected probabilities.**

**Benford's Law detection** offers a statistical approach to identifying potential data manipulation or errors in financial records. The `xbtlin/ai-berkshire` repository provides a lightweight, zero-dependency implementation in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) that requires no external packages beyond Python's standard library.

## How the Benford's Law Detector Works

The `benford_check` function (defined at lines 14-81 of [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py)) implements a complete statistical pipeline for first-digit analysis.

### Pre-computed Benford Probabilities

The module stores expected first-digit frequencies in the `_BENFORD` dictionary (lines 11-12). This dictionary maps digits 1-9 to their theoretical logarithmic probabilities using the formula `math.log10(1 + 1/d)`:

```python
_BENFORD = {d: math.log10(1 + 1/d) for d in range(1, 10)}

```

These values represent the expected frequency for each leading digit according to Benford's Law.

### Leading-Digit Extraction and Validation

For each input value, the function converts the number to a positive float, removes the order of magnitude, and isolates the leading digit (lines 22-28). This normalization ensures that `1234`, `12.34`, and `0.001234` all resolve to the same leading digit `1`.

### Sample-Size Guards

Benford's Law requires sufficient data volume to produce reliable results. The implementation enforces a minimum threshold of **50 non-zero observations** (lines 30-34). If the sample size falls below this threshold, the function prints a warning and adjusts its confidence accordingly.

### Statistical Metrics

The detector calculates two key goodness-of-fit metrics:

1. **MAD (Mean Absolute Deviation)** – Computed as the average absolute difference between observed and expected frequencies (lines 42-43). This metric provides a straightforward measure of deviation from Benford's distribution.

2. **Chi-square** – Calculated at lines 45-46 to provide an additional statistical significance test comparing observed versus expected distributions.

### Conformity Classification

Based on the MAD calculation, the function classifies results into four categories (lines 48-55):

- **Close (高度符合)** – MAD < 0.006
- **Acceptable (可接受)** – MAD between 0.006 and 0.012
- **Marginally Acceptable (边缘)** – MAD between 0.012 and 0.015
- **Non-conforming (不符合)** – MAD ≥ 0.015

The function returns a dictionary containing `mad`, `chi2`, `conformity`, and `is_conforming` boolean flag (lines 60-81), while simultaneously printing a formatted table of observed versus expected frequencies.

## Running Benford's Law Detection in financial_rigor.py

### Command-Line Interface

Invoke the detector directly from the terminal using the built-in CLI:

```bash
python3 tools/financial_rigor.py benford \
    --values '[1234, 5678, 9012, 3456, 7890, 2345, 6789, 1234, 5678, 9012]'

```

For reliable results, provide at least 50 numeric values in the JSON array. The CLI outputs both the comparison table and the final conformity verdict.

### Programmatic Usage

Import the function directly into your analysis scripts:

```python
from tools.financial_rigor import benford_check

# Ensure at least 50 non-zero values for statistical validity

financial_data = [1234, 5678, 9012, 3456, 7890, 2345] * 10
result = benford_check(financial_data)

print(f"MAD: {result['mad']}")
print(f"Conformity: {result['conformity']}")
print(f"Conforming: {result['is_conforming']}")

```

The function returns structured data suitable for automated reporting pipelines or dashboard integration.

## Interpreting Anomaly Detection Results

When reviewing `benford_check` output, focus on the **MAD value** and **conformity classification**:

- **MAD < 0.006** indicates the dataset follows Benford's distribution closely, suggesting natural financial data without obvious manipulation.

- **MAD ≥ 0.015** triggers a "Non-conforming" classification, serving as a red flag requiring investigation into potential adjustments, errors, or fraudulent entries.

- **Chi-square values** provide secondary confirmation; high values combined with elevated MAD strengthen the case for anomalous data.

Treat "Marginally Acceptable" results as indicators to expand your sample size or cross-reference with additional validation methods available in the financial rigor toolkit.

## Summary

- The `benford_check` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) provides zero-dependency Benford's Law analysis for financial anomaly detection.
- The implementation requires **≥50 non-zero observations** for reliable statistical output (lines 30-34).
- **MAD thresholds** classify results as Close (<0.006), Acceptable (0.006-0.012), Marginally Acceptable (0.012-0.015), or Non-conforming (≥0.015).
- The function returns a dictionary with `mad`, `chi2`, `conformity`, and `is_conforming` keys for programmatic processing.
- Both CLI and Python API interfaces are available, requiring only the Python standard library.

## Frequently Asked Questions

### What is the minimum sample size required for Benford's law detection?

The `benford_check` function requires **at least 50 non-zero observations** to produce reliable results. According to the source code at lines 30-34, samples below this threshold trigger a warning message indicating insufficient data volume for accurate Benford analysis.

### What MAD threshold indicates potential data manipulation?

A **MAD (Mean Absolute Deviation) value ≥ 0.015** results in a "Non-conforming" classification, indicating significant deviation from Benford's expected distribution. As implemented in lines 48-55 of [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py), this threshold serves as the primary red flag for potential data manipulation, errors, or artificial adjustments in financial records.

### Can the detector process negative numbers or zero values?

The implementation automatically converts values to positive floats during leading-digit extraction (lines 22-28), allowing it to handle negative numbers. However, **zero values are excluded** from the analysis since they have no first digit. The sample-size check specifically counts "non-zero observations" to ensure sufficient valid data points for statistical testing.

### How does this compare to other validation tools in the ai-berkshire repository?

While [`financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/financial_rigor.py) contains multiple validation utilities, the `benford_check` function uniquely provides **distribution-based anomaly detection** without requiring external dependencies. Unlike rule-based validators, it detects subtle patterns that deviate from expected natural distributions, making it complementary to the other financial validation tools referenced in [`AGENTS.md`](https://github.com/xbtlin/ai-berkshire/blob/main/AGENTS.md) and integrated through [`scripts/sync-codex-skills.py`](https://github.com/xbtlin/ai-berkshire/blob/main/scripts/sync-codex-skills.py).