# Benford's Law Detection in financial_rigor.py: Identifying Financial Data Anomalies

> Discover how financial data anomalies are detected using Benford's Law in financial_rigor.py. Identify nonconforming datasets and potential manipulation with statistical analysis.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-26

---

**The `benford_check` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) detects anomalies by measuring how closely the leading digits of financial values follow Benford's expected distribution, flagging "Nonconforming" datasets when statistical deviations exceed thresholds that suggest data manipulation or rounding artifacts.**

The `ai-berkshire` repository provides a lightweight Benford's Law audit tool in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) designed for on-the-fly validation of financial series. The `benford_check` function compares empirical leading-digit frequencies against the theoretical `log10(1 + 1/d)` distribution to identify statistical irregularities that warrant deeper investigation.

## How `benford_check` Detects Anomalies

The detection workflow follows a rigorous statistical pipeline implemented between lines 14 and 81 of [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py). Each step is optimized for quick validation of numeric financial series.

### Leading Digit Extraction and Sample Validation

For each positive number in the input list, the function isolates the leading digit using logarithmic manipulation. According to lines 20-28, the code calculates `sig = 10 ** (log10(v) - floor(log10(v)))` and casts the result to an integer, yielding digits 1 through 9.

Before proceeding with statistical analysis, the function enforces a sample-size guard at lines 30-33. If fewer than 50 digits are collected, the function returns `None`, because Benford's Law requires sufficient volume for the expected distribution to emerge meaningfully.

### Statistical Conformance Metrics

The function computes two primary conformance metrics against the pre-computed expected frequencies stored in the `_BENFORD` constant at line 11 (calculated as `log10(1 + 1/d)` for digits 1-9).

First, it tallies occurrences and normalizes them to empirical frequencies (lines 35-40). Then it calculates:

- **Mean Absolute Deviation (MAD)** at line 42 – the average absolute gap between observed and expected frequencies
- **Chi-square statistic** at lines 44-45 – a classic goodness-of-fit measure

### Anomaly Classification Thresholds

Lines 48-55 implement a four-tier classification system based on MAD values:

- **MAD < 0.006**: "Close" conformity
- **MAD < 0.012**: "Acceptable" conformity  
- **MAD < 0.015**: "Marginally Acceptable"
- **MAD ≥ 0.015**: "Nonconforming" – the primary anomaly flag

When reporting (lines 63-71), the function prints a diagnostic table flagging individual digits whose deviation exceeds ±0.03. The final decision (lines 74-79) returns a boolean `is_conforming` value (`True` only when `mad < 0.015`), along with the specific conformity classification string.

## Implementation Details and Code Usage

The `benford_check` function is accessible both as a CLI tool and as a Python module, ensuring consistent anomaly detection across interactive scripts and automated pipelines.

### Command-Line Interface

Invoke the audit directly from the terminal for quick validation of financial datasets:

```bash
python3 tools/financial_rigor.py benford \
    --values '[12345, 67890, 23456, 34567, 45678, 56789, 67890, ...]'

```

The script outputs a diagnostic table showing observed versus expected percentages, the calculated MAD and chi-square values, and a clear pass/fail conformity status.

### Programmatic Integration in Python

Import the function to validate financial series within automated workflows:

```python
from tools.financial_rigor import benford_check

# Example: revenue figures in millions

revenues = [
    1234, 2345, 3456, 4567, 5678, 6789, 7890,
    8912, 9123, 10123, 11234, 12345, 13456,
    # ... (ensure >50 entries for reliable results)

]

report = benford_check(revenues)

print(report["conformity"])      # e.g., "Acceptable (可接受)"

print(report["is_conforming"])   # True or False

print(report["mad"])             # 0.0081

print(report["chi2"])            # 12.34

```

The return payload at line 81 contains `mad`, `chi2`, `conformity` text, and the `is_conforming` boolean, enabling downstream logic to trigger alerts when nonconforming data is detected.

## Types of Anomalies Identified

The `benford_check` function specifically identifies **distributional anomalies** in the leading-digit frequency of financial datasets:

- **Systematic deviations from Benford's Law**: When the empirical distribution of first digits (1-9) significantly diverges from the expected logarithmic curve, typically indicating artificial data construction or excessive rounding.
- **Nonconforming classifications**: Datasets with MAD ≥ 0.015 that fail the statistical goodness-of-fit test, signaling potential data manipulation or transcription errors.
- **Per-digit outliers**: Individual digits showing frequency deviations greater than ±0.03 from expected values, which may reveal specific biases (such as excessive use of the digit 5 in fabricated numbers).

These statistical red flags do not prove fraud but highlight datasets requiring manual review, particularly useful when screening large volumes of financial figures in the AI-Berkshire investment workflow.

## Summary

- **`benford_check`** in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) validates financial data conformity to Benford's Law through statistical analysis of leading-digit distributions.
- The function extracts leading digits using logarithmic calculation (lines 20-28) and requires a minimum sample size of 50 values to proceed.
- Conformance is quantified using **Mean Absolute Deviation** (MAD) and **chi-square** statistics compared against the theoretical `_BENFORD` distribution (line 11).
- **MAD thresholds** classify results as Close (<0.006), Acceptable (<0.012), Marginally Acceptable (<0.015), or Nonconforming (≥0.015).
- A **"Nonconforming"** flag indicates anomalous data that deviates significantly from expected Benford frequencies, suggesting potential manipulation or data quality issues.

## Frequently Asked Questions

### What is the minimum sample size for reliable Benford's Law detection in financial_rigor.py?

The `benford_check` function enforces a strict minimum of **50 digits** before performing statistical analysis. According to lines 30-33 in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py), the function returns `None` if the input contains fewer than 50 values, as smaller samples lack the statistical power to reliably approximate the expected Benford distribution.

### What MAD threshold indicates nonconforming data in benford_check?

The function classifies data as **"Nonconforming"** when the Mean Absolute Deviation (MAD) is **0.015 or greater**. This threshold is implemented at lines 48-55 in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py). Values below this threshold receive classifications of "Close," "Acceptable," or "Marginally Acceptable" depending on specific MAD ranges, while the `is_conforming` boolean returns `False` only at or above the 0.015 threshold.

### How does benford_check extract leading digits from financial values?

The function extracts leading digits using a logarithmic transformation implemented at lines 20-28. For each positive number `v`, it calculates `sig = 10 ** (log10(v) - floor(log10(v)))` to isolate the significand, then casts this to an integer. This mathematical approach efficiently yields the first digit (1-9) without string manipulation, handling arbitrary numeric scales automatically.

### Can benford_check detect specific types of financial fraud?

The function detects **distributional anomalies** that correlate with certain fraud patterns, such as excessive rounding, made-up numbers that avoid specific digits, or manipulated figures that cluster unnaturally. However, it cannot identify the *mechanism* of manipulation—only that the data deviates from Benford's Law. A "Nonconforming" result (MAD ≥ 0.015) or per-digit deviations exceeding ±0.03 serve as red flags requiring manual investigation, but legitimate business factors (like price caps at $99) can also cause deviations.