# How to Use the report_audit.py Tool to Verify Research Report Quality

> Verify research report quality with report_audit.py. This tool extracts numeric data, samples it, and validates against independent sources for a PASS/FAIL verdict. Ensure accuracy easily.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-28

---

**The report_audit.py command-line utility extracts numeric data points from Markdown research reports, randomly samples approximately 15% of them, and validates the sampled values against independent sources to generate a PASS/FAIL quality verdict.**

The [`report_audit.py`](https://github.com/xbtlin/ai-berkshire/blob/main/report_audit.py) tool in the [xbtlin/ai-berkshire](https://github.com/xbtlin/ai-berkshire) repository provides a dependency-free Python solution for auditing financial research reports. It parses Markdown files to identify numerical claims—including percentages, Chinese units like "亿" (hundred million), and multiples—then statistically verifies their accuracy against external data providers such as macrotrends or stockanalysis.

## The Three-Stage Audit Workflow

The tool implements a rigorous pipeline defined in [`tools/report_audit.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/report_audit.py) comprising three distinct stages.

### Stage 1: Data Point Extraction

The `extract_data_points()` function parses the input Markdown file to locate numeric values using comprehensive regex patterns. It recognizes percentages, table cells, key-value lines, and Chinese financial notation including "亿" and "万亿". The extraction engine handles ASCII `+/-`, Unicode minus (U+2212), en-dash (U+2013), and full-width minus (U+FF0D) through the `_SIGN` regex and `_clean_num()` implementation to ensure negative values like `-1.72%` are never interpreted as positive.

### Stage 2: Random Sampling

The `sample_points()` function performs statistical subset selection, defaulting to 15% of extracted data points while enforcing hard bounds of minimum 3 and maximum 30 samples. This balanced approach ensures manageable verification workloads without sacrificing statistical relevance. Users can adjust the sampling rate with `--ratio` or enforce reproducibility via `--seed`.

### Stage 3: Verdict Generation

The `render_verdict()` function computes relative deviations between reported values and fetched reference data, assigning PASS, WARNING, or FAIL status to each point. It outputs a colored console report and returns exit code `0` for overall PASS or `1` for FAIL, enabling seamless integration with CI/CD pipelines.

## Step-by-Step Usage Guide

### Extract and Sample Data Points

Initiate the audit workflow by extracting data points from your research report:

```bash
python3 tools/report_audit.py extract \
  --report reports/your-report-2023-09-01.md \
  --ratio 0.15 \
  --seed 42 \
  --dry-run

```

The `--dry-run` flag prints a human-readable table of sampled points without emitting JSON, useful for preliminary review. Omit this flag to generate the JSON template required for verification. The command displays line numbers, labels, reported values, and units for each sampled data point.

### Populate the JSON Template

The extraction step produces a JSON array with placeholders for verification:

```json
{
  "id": 1,
  "label": "营业收入",
  "reported_value": 7518,
  "unit": "亿",
  "line_number": 12,
  "raw_text": "...",
  "fetched_value": null,
  "fetched_source": "",
  "fetched_value2": null,
  "fetched_source2": ""
}

```

Manually populate `fetched_value` with corresponding data from authoritative sources such as `macrotrends.net`, `stockanalysis.com`, `aastocks.com`, or `eastmoney.com`. Optionally provide `fetched_value2` for secondary source validation. The tool intentionally requires manual data entry to maintain zero dependencies while allowing flexible integration with any data provider.

### Run the Verdict Command

Execute the verification against your populated data:

```bash
python3 tools/report_audit.py verdict \
  --results "$(cat audit.json)" \
  --report "Your Report Title" \
  --output-json

```

The `--output-json` option emits machine-readable results alongside the console summary. The command calculates percentage deviations and categorizes discrepancies as warnings (minor variance) or failures (significant error).

## Practical Example: Complete Workflow

The following demonstrates a full audit cycle for a Tencent research report:

```bash

# Extract and sample with reproducible seed

python3 tools/report_audit.py extract \
  --report reports/腾讯/腾讯-research-20260408.md \
  --seed 7 > audit.json

# Edit audit.json to populate fetched values from external sources

# Generate final quality assessment

python3 tools/report_audit.py verdict \
  --results "$(cat audit.json)" \
  --report "腾讯 2026-04-08 Research"

```

The console displays individual assessments with deviation percentages:

```

✅ 通过 [ 1] 营业收入 · 合计           7518.00 亿  → macrotrends: 7518.00 (偏差 0.00%)
⚠️  警告 [ 2] 毛利率                  45.20 %   → stockanalysis: 45.10 (偏差 0.22%)
❌ 不通过 [ 3] 负债率                  120.00 % → macrotrends: 115.00 (偏差 4.35%)

```

## Robustness Mechanisms

The implementation includes specific safeguards for production reliability documented in the source code.

### Encoding Resilience

The `_force_utf8_stdio()` function forces UTF-8 encoding for stdout and stderr streams, preventing `UnicodeEncodeError` exceptions on Windows systems using GBK console encoding.

### Value Sanity Filtering

The extraction logic discards absurd values where `abs(val) > 1e15` while correctly preserving legitimate negative large numbers. This filtering prevents corrupted data or parsing artifacts from contaminating the audit results.

### Comprehensive Testing

The [`tests/test_report_audit.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tests/test_report_audit.py) suite validates negative-sign extraction across Unicode variants, GBK console survival, and proper handling of extreme negative values, ensuring the tool behaves correctly across diverse runtime environments.

## Summary

- **Three-stage verification**: Extraction via `extract_data_points()`, sampling via `sample_points()`, and validation via `render_verdict()` in [`tools/report_audit.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/report_audit.py).
- **Zero dependencies**: Requires only Python standard library; no external packages needed.
- **Unicode-aware parsing**: Correctly handles multiple minus sign representations (U+002D, U+2212, U+2013, U+FF0D) and Chinese financial units.
- **Configurable sampling**: Default 15% ratio with 3-30 sample bounds, controlled via `--ratio` and `--seed` parameters.
- **CI/CD integration**: Returns exit code 0 for PASS and 1 for FAIL, with `--output-json` for automated pipelines.
- **Defensive programming**: UTF-8 encoding enforcement and value bounds checking prevent runtime errors.

## Frequently Asked Questions

### What file formats does report_audit.py support?

The tool specifically processes Markdown (`.md`) files. It scans text content, tables, and key-value lines to identify numeric patterns, with particular optimization for financial reports containing Chinese characters and units like "亿" or "万亿". No other file formats are supported.

### How does the sampling algorithm ensure critical data is verified?

The `sample_points()` function implements random sampling with configurable seeds for reproducibility. While the default 15% ratio provides statistical efficiency, the mandatory minimum of 3 samples guarantees baseline coverage even for short reports. For high-stakes research, users can increase the `--ratio` parameter to sample more points or use `--dry-run` to review exactly which data points will be audited before committing to verification.

### Can the data fetching process be automated?

The current architecture requires manual population of the `fetched_value` fields. The tool deliberately separates data sourcing from audit logic to maintain zero dependencies and allow integration with any proprietary or public data source. Users can create wrapper scripts to automate fetching from specific APIs (macrotrends, stockanalysis, etc.) while using [`report_audit.py`](https://github.com/xbtlin/ai-berkshire/blob/main/report_audit.py) solely for the verification engine.

### Why does the tool discard values larger than 1e15?

The bound check `abs(val) > 1e15` serves as a sanity filter against numeric overflow, corrupted table cells, or copy-paste artifacts that might appear in Markdown content. This threshold captures legitimate large-scale financial figures while eliminating mathematical anomalies that would otherwise skew deviation calculations in the verdict stage.