# How Confidence Is Handled in Private Company Research with Sparse Financial Data

> Learn how AI-Berkshire handles confidence in private company research with sparse data. Discover confidence-weighted data and cross-validation techniques.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: deep-dive
- Published: 2026-07-26

---

**The AI-Berkshire framework treats every data point as a confidence-weighted datum, assigning 🟢 High, 🟡 Medium, or 🔴 Low tags based on source reliability and requiring cross-validation from two independent sources before aggregation.**

The AI-Berkshire open-source framework addresses the fundamental challenge of evaluating unlisted companies by implementing a rigorous confidence-weighting system. When analyzing private firms with incomplete financial statements or missing regulatory filings, the system explicitly propagates uncertainty through every stage of the research pipeline rather than fabricating certainty. This methodology ensures that sparse financial data never masquerades as ground truth.

## Confidence Tagging and Source Classification

### The Three-Level Confidence System

Every metric collected during private company research carries a visual confidence tag directly in the Markdown output. The framework uses a simple emoji-based system defined in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md):

- **🟢 High**: Audited financials, SEC filings, official disclosures, and annual reports
- **🟡 Medium**: Reputable business journalism (Bloomberg, Reuters, 36Kr, The Information) and industry research reports
- **🔴 Low**: Blogs, social media posts, rumors, and employee leaks

These tags appear immediately adjacent to values in research tables, such as `Revenue $120 M 🟢` or `EBITDA $12 M (估算) 🔴`, making uncertainty transparent to the reader.

### Source-Based Confidence Assignment

The confidence level is not guessed but derived from a predefined **data source matrix** (`数据来源矩阵`) located in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md). When the six parallel agents—`business-decoder`, `financial-detective`, `competitive-mapper`, `risk-governance-analyst`, `tech-ip-analyst`, and `signal-miner`—ingest data, they automatically apply the matrix rules:

- **Official filings** (`prospectus`, `年报`, `官方披露`) → 🟢
- **Verified journalism** (`Bloomberg`, `Reuters`, `行业研报`) → 🟡  
- **Unverified sources** (`博客`, `传闻`, `社交媒体`, `员工爆料`) → 🔴

This enforcement happens at ingestion, ensuring downstream components receive already-tagged data.

## Information Richness Ratings

Beyond individual data points, the framework assigns an **Information Richness Rating** (A/B/C) documented in [`README_EN.md`](https://github.com/xbtlin/ai-berkshire/blob/main/README_EN.md). This coarse rating reminds analysts that "more data ≠ more certainty" and influences how much weight is given to aggregate scores during final synthesis. A company with extensive but low-confidence filings might receive a lower composite reliability score than one with sparse but audited documents.

## Cross-Validation and Data Integrity

The framework mandates **dual-source verification** for all key metrics. According to the "财务数据交叉验证" section in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md), each critical figure must be corroborated by at least two independent sources.

When sources disagree, the system does not auto-correct or hide the discrepancy. Instead, it retains all values with their respective confidence tags and appends a reconciliation note explaining the variance. This preserves the full uncertainty spectrum for human reviewers.

## Aggregating Confidence Across Agents

Final valuation ranges incorporate a **composite confidence** score calculated by the `financial-detective` agent. As detailed in the "估值综合判断" section of [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md), this composite derives from a weighted average of individual data point confidences.

Numeric mappings (High=3, Medium=2, Low=1) allow mathematical aggregation. If the weighted average exceeds 2.5, the final judgment receives a 🟢; between 1.5 and 2.5 receives 🟡; below 1.5 receives 🔴. This score appears alongside each valuation method in the final output table.

## Explicit Handling of Missing Data

When no reliable source exists for a requested metric, the framework follows the "应对原则" (Response Principles) in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md): the field is left blank with an explicit annotation **"数据缺失"** (data missing) rather than interpolating or fabricating a number. This prevents hallucinated figures from contaminating valuation models.

## Implementation Examples

Below are concrete implementations showing how confidence tagging works in practice.

### Markdown Output Format

Research reports use standard Markdown tables with dedicated confidence columns:

```markdown
| 指标 | 数据 | 来源 | 置信度 |
|------|------|------|--------|
| 2024 Q1 Revenue | $85 M | Bloomberg article "Company X Q1 Results" | 🟢 |
| 2024 Q1 EBITDA | $12 M (估算) | 业务报告推算 | 🔴 |

```

### Source-to-Confidence Mapping

This Python helper implements the source matrix logic:

```python
def confidence_for_source(source: str) -> str:
    """Return the confidence emoji for a given source descriptor."""
    high = {"prospectus", "SEC filing", "官方披露", "年报"}
    medium = {"Bloomberg", "Reuters", "36氪", "The Information", "行业研报"}
    low = {"博客", "传闻", "社交媒体", "员工爆料"}

    if source in high:
        return "🟢"
    if source in medium:
        return "🟡"
    if source in low:
        return "🔴"
    return "🔴"  # default to low if unknown

```

### Composite Confidence Calculation

The aggregation logic used by the `financial-detective` agent:

```python
def aggregate_confidence(tags: list[float]) -> str:
    """Convert a list of numeric confidence values (1-3) to an overall emoji."""
    avg = sum(tags) / len(tags)
    if avg >= 2.5:
        return "🟢"
    if avg >= 1.5:
        return "🟡"
    return "🔴"

# Example usage:

tags = [3, 2, 1]          # 🟢, 🟡, 🔴 → numeric 3,2,1

print(aggregate_confidence(tags))   # → 🟡

```

## Summary

- **Every data point** in AI-Berkshire carries an explicit confidence tag (🟢/🟡/🔴) based on the source type matrix in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md).
- **Dual-source validation** is required before any metric enters the final report, with disagreements preserved rather than resolved.
- **Composite confidence** is mathematically derived from the weighted average of individual tags by the `financial-detective` agent.
- **Missing data** is explicitly marked as "数据缺失" rather than estimated, preventing fabricated certainty.
- **Information Richness Ratings** (A/B/C) provide a meta-layer of certainty assessment across the entire research project.

## Frequently Asked Questions

### How does the AI-Berkshire framework assign confidence levels to different data sources?

The framework uses a predefined **data source matrix** located in the "数据来源矩阵" section of [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md). This matrix categorizes inputs into three tiers: official regulatory filings and audited prospectuses receive 🟢 (High), established financial journalism and industry reports receive 🟡 (Medium), and unverified sources like blogs or social media receive 🔴 (Low).

### What happens when two sources disagree on a financial metric?

According to the "财务数据交叉验证" section, the system retains **all conflicting values** with their respective confidence tags and appends a reconciliation note explaining the discrepancy. Rather than forcing consensus, the framework presents the uncertainty to the user, who can then judge which source merits more trust based on the confidence indicators.

### How is composite confidence calculated for final valuation scores?

The `financial-detective` agent converts emoji tags to numeric values (🟢=3, 🟡=2, 🔴=1) and computes a weighted average across all inputs for a given valuation method. As implemented in the "估值综合判断" logic, averages ≥2.5 receive a final 🟢, ≥1.5 receive 🟡, and lower values receive 🔴, clearly signaling the reliability of the resulting valuation range.

### What does the system do when no reliable financial data is available?

Following the "应对原则" in [`skills/private-company-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/private-company-research.md), the framework leaves the field explicitly blank with the annotation **"数据缺失"** (data missing). This prevents the LLM agents from hallucinating plausible-sounding but fabricated numbers, ensuring that investment decisions are never based on synthetic data.