# How the earnings-review Skill Parses Primary Sources Like 10-K Filings in ai-berkshire

> Learn how the earnings-review skill parses SEC 10-K filings from raw HTML and PDF into structured JSON for LLM analysis using Python libraries.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The earnings-review skill fetches raw HTML and PDF filings directly from SEC EDGAR and Taiwan MOPS, then parses them into structured JSON sections using BeautifulSoup, pdfminer.six, and regex-based heading extraction, allowing the LLM to answer questions solely from primary source material.**

The **earnings-review** skill in the open-source **xbtlin/ai-berkshire** repository eliminates reliance on secondary analyst summaries by programmatically downloading and structuring data directly from regulatory filings. Instead of reading pre-processed databases, the skill handles **10-K** reports from the U.S. Securities and Exchange Commission, **IO-K** filings from Chinese markets, and Taiwan Market Observation Post System (MOPS) documents, converting unstructured HTML and PDFs into machine-readable sections that large language models can reference with source-level accuracy.

## Fetching Raw Filings from SEC EDGAR and MOPS

The ingestion pipeline starts in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py), where the `fetch_sec_filing()` function constructs direct URLs to the SEC EDGAR archive.

```python
def fetch_sec_filing(ticker: str, year: int, form: str = "10-K") -> str:
    """
    Pull the HTML version of the filing from SEC EDGAR.
    Returns the raw HTML as a string.
    """
    base = "https://www.sec.gov/Archives/edgar/data"
    # Resolves CIK via local ticker-CIK map, then constructs:

    # https://www.sec.gov/Archives/edgar/data/<CIK>/<accession-number>/index.html

```

For Taiwanese equities, the companion function `fetch_tw_filing()` interfaces with the Taiwan MOPS API to retrieve PDF documents, ensuring coverage across jurisdictions without intermediaries.

## Converting Raw Documents to Plain Text

Once downloaded, filings undergo extraction via two distinct paths depending on format. For HTML documents, the `html_to_text()` function uses **BeautifulSoup** to strip navigation elements and normalize whitespace.

```python
from bs4 import BeautifulSoup

def html_to_text(html: str) -> str:
    soup = BeautifulSoup(html, "html.parser")
    # Remove navigation, scripts, style tags

    for s in soup(["script", "style", "nav"]):
        s.decompose()
    return soup.get_text(separator="\n")

```

For PDF filings—common in non-SEC jurisdictions—the skill utilizes **pdfminer.six** through a `pdf_to_text()` helper that performs equivalent cleanup and text extraction.

## Section Extraction via Regular Expressions

The critical parsing logic relies on SEC filings following a standardized "Item X" heading structure (e.g., "Item 1. Business", "Item 1A. Risk Factors", "Item 7. Management’s Discussion and Analysis"). The `split_into_sections()` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) applies a compiled regex pattern to locate these boundaries.

```python
import re
from typing import Dict

SECTION_RE = re.compile(r"Item\s+(\d+[A-Z]?)\.\s+([^\n]+)", re.IGNORECASE)

def split_into_sections(text: str) -> Dict[str, str]:
    sections = {}
    matches = list(SECTION_RE.finditer(text))
    for i, m in enumerate(matches):
        start = m.end()
        end = matches[i+1].start() if i+1 < len(matches) else len(text)
        title = f"Item {m.group(1)} – {m.group(2).strip()}"
        sections[title] = text[start:end].strip()
    return sections

```

This regex captures the item number (including letter suffixes like 1A) and the descriptive text, partitioning the document into discrete sections that preserve the original source hierarchy.

## Structuring Data for LLM Consumption

The parsed sections are normalized into a unified JSON structure defined in [`skills/earnings-review.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/earnings-review.md). This payload includes metadata and the extracted section text, which the skill passes to the LLM with explicit instructions to reference only the provided primary material.

```json
{
  "ticker": "AAPL",
  "year": 2025,
  "form": "10-K",
  "sections": {
    "Item 1 – Business": "...",
    "Item 1A – Risk Factors": "...",
    "Item 7 – MD&A": "...",
    "Item 8 – Financial Statements": "..."
  }
}

```

The prompt template in [`codex-skills/earnings-review/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-skills/earnings-review/SKILL.md) instructs the model: *"Use the following parsed filing sections to answer the user's question. Do not rely on any secondary analyst reports."* This constraint grounds the LLM's responses in the actual filing text.

## Complete Implementation Example

The following runnable Python example demonstrates the full pipeline from fetch to structured sections:

```python
from tools.financial_rigor import fetch_sec_filing, html_to_text, split_into_sections

# Download Apple's 2025 10-K filing

html_content = fetch_sec_filing("AAPL", 2025, form="10-K")

# Convert HTML to clean text

plain_text = html_to_text(html_content)

# Extract sections using Item heading regex

sections = split_into_sections(plain_text)

# Access specific primary source sections

print(sections.get("Item 1 – Business", "Section not found"))
print(sections.get("Item 1A – Risk Factors", "Section not found"))

```

For command-line usage, the repository provides a wrapper that automates this pipeline:

```bash
ai-berkshire --skill earnings-review --ticker TSLA --year 2025

```

This CLI tool assembles the JSON payload, sends it to the LLM, and returns formatted answers citing only the extracted primary source sections.

## Summary

- **Direct Retrieval**: The skill downloads raw filings from official repositories (SEC EDGAR for U.S. equities, Taiwan MOPS for Taiwanese stocks) rather than commercial data feeds.
- **Format Agnostic**: Supports both HTML (BeautifulSoup) and PDF (pdfminer.six) extraction pipelines within [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py).
- **Structured Sectioning**: Uses regex pattern `Item\s+(\d+[A-Z]?)\.\s+([^\n]+)` to split documents into standardized headings like "Item 1 – Business".
- **LLM Integration**: Outputs normalized JSON consumed by the prompt template in [`skills/earnings-review.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/earnings-review.md), enforcing source-grounded responses without secondary interpretation.

## Frequently Asked Questions

### What file types does the earnings-review skill support?

The skill processes both **HTML** and **PDF** filings. For SEC EDGAR 10-K and 10-Q reports, it typically parses HTML versions using `html_to_text()`. For Taiwanese MOPS filings and other jurisdictions that provide PDFs only, it utilizes `pdf_to_text()` with pdfminer.six to extract content before applying the section-splitting regex.

### How does the skill handle different filing formats across countries?

The `split_into_sections()` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) accepts the plain text output regardless of original format (HTML or PDF). While the regex pattern targets SEC-style "Item X" headings, the same logic applies to IO-K and other regional filings by mapping their equivalent heading structures to the same extraction logic, making the skill extensible to new jurisdictions by adjusting the regex or adding jurisdiction-specific heading mappings.

### Can I use this parser independently of the LLM skill?

Yes. The core parsing functions—`fetch_sec_filing()`, `html_to_text()`, and `split_into_sections()`—are importable standalone utilities within [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py). You can use these to build financial data pipelines, perform bulk extraction of specific sections like "Risk Factors" across multiple years, or integrate the structured output into quantitative analysis workflows without invoking the LLM components.

### Where is the prompt template that prevents the LLM from using secondary sources?

The constraint instructing the LLM to use only primary source material resides in [`skills/earnings-review.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/earnings-review.md) and is duplicated in the auto-generated [`codex-skills/earnings-review/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-skills/earnings-review/SKILL.md). These files contain the system prompt that accompanies the structured JSON payload, explicitly forbidding the model from referencing analyst reports, news articles, or internal training data when answering questions about the filing.