How the earnings-review Skill Parses Primary Sources Like 10-K Filings in ai-berkshire
The earnings-review skill fetches raw HTML and PDF filings directly from SEC EDGAR and Taiwan MOPS, then parses them into structured JSON sections using BeautifulSoup, pdfminer.six, and regex-based heading extraction, allowing the LLM to answer questions solely from primary source material.
The earnings-review skill in the open-source xbtlin/ai-berkshire repository eliminates reliance on secondary analyst summaries by programmatically downloading and structuring data directly from regulatory filings. Instead of reading pre-processed databases, the skill handles 10-K reports from the U.S. Securities and Exchange Commission, IO-K filings from Chinese markets, and Taiwan Market Observation Post System (MOPS) documents, converting unstructured HTML and PDFs into machine-readable sections that large language models can reference with source-level accuracy.
Fetching Raw Filings from SEC EDGAR and MOPS
The ingestion pipeline starts in tools/financial_rigor.py, where the fetch_sec_filing() function constructs direct URLs to the SEC EDGAR archive.
def fetch_sec_filing(ticker: str, year: int, form: str = "10-K") -> str:
"""
Pull the HTML version of the filing from SEC EDGAR.
Returns the raw HTML as a string.
"""
base = "https://www.sec.gov/Archives/edgar/data"
# Resolves CIK via local ticker-CIK map, then constructs:
# https://www.sec.gov/Archives/edgar/data/<CIK>/<accession-number>/index.html
For Taiwanese equities, the companion function fetch_tw_filing() interfaces with the Taiwan MOPS API to retrieve PDF documents, ensuring coverage across jurisdictions without intermediaries.
Converting Raw Documents to Plain Text
Once downloaded, filings undergo extraction via two distinct paths depending on format. For HTML documents, the html_to_text() function uses BeautifulSoup to strip navigation elements and normalize whitespace.
from bs4 import BeautifulSoup
def html_to_text(html: str) -> str:
soup = BeautifulSoup(html, "html.parser")
# Remove navigation, scripts, style tags
for s in soup(["script", "style", "nav"]):
s.decompose()
return soup.get_text(separator="\n")
For PDF filings—common in non-SEC jurisdictions—the skill utilizes pdfminer.six through a pdf_to_text() helper that performs equivalent cleanup and text extraction.
Section Extraction via Regular Expressions
The critical parsing logic relies on SEC filings following a standardized "Item X" heading structure (e.g., "Item 1. Business", "Item 1A. Risk Factors", "Item 7. Management’s Discussion and Analysis"). The split_into_sections() function in tools/financial_rigor.py applies a compiled regex pattern to locate these boundaries.
import re
from typing import Dict
SECTION_RE = re.compile(r"Item\s+(\d+[A-Z]?)\.\s+([^\n]+)", re.IGNORECASE)
def split_into_sections(text: str) -> Dict[str, str]:
sections = {}
matches = list(SECTION_RE.finditer(text))
for i, m in enumerate(matches):
start = m.end()
end = matches[i+1].start() if i+1 < len(matches) else len(text)
title = f"Item {m.group(1)} – {m.group(2).strip()}"
sections[title] = text[start:end].strip()
return sections
This regex captures the item number (including letter suffixes like 1A) and the descriptive text, partitioning the document into discrete sections that preserve the original source hierarchy.
Structuring Data for LLM Consumption
The parsed sections are normalized into a unified JSON structure defined in skills/earnings-review.md. This payload includes metadata and the extracted section text, which the skill passes to the LLM with explicit instructions to reference only the provided primary material.
{
"ticker": "AAPL",
"year": 2025,
"form": "10-K",
"sections": {
"Item 1 – Business": "...",
"Item 1A – Risk Factors": "...",
"Item 7 – MD&A": "...",
"Item 8 – Financial Statements": "..."
}
}
The prompt template in codex-skills/earnings-review/SKILL.md instructs the model: "Use the following parsed filing sections to answer the user's question. Do not rely on any secondary analyst reports." This constraint grounds the LLM's responses in the actual filing text.
Complete Implementation Example
The following runnable Python example demonstrates the full pipeline from fetch to structured sections:
from tools.financial_rigor import fetch_sec_filing, html_to_text, split_into_sections
# Download Apple's 2025 10-K filing
html_content = fetch_sec_filing("AAPL", 2025, form="10-K")
# Convert HTML to clean text
plain_text = html_to_text(html_content)
# Extract sections using Item heading regex
sections = split_into_sections(plain_text)
# Access specific primary source sections
print(sections.get("Item 1 – Business", "Section not found"))
print(sections.get("Item 1A – Risk Factors", "Section not found"))
For command-line usage, the repository provides a wrapper that automates this pipeline:
ai-berkshire --skill earnings-review --ticker TSLA --year 2025
This CLI tool assembles the JSON payload, sends it to the LLM, and returns formatted answers citing only the extracted primary source sections.
Summary
- Direct Retrieval: The skill downloads raw filings from official repositories (SEC EDGAR for U.S. equities, Taiwan MOPS for Taiwanese stocks) rather than commercial data feeds.
- Format Agnostic: Supports both HTML (BeautifulSoup) and PDF (pdfminer.six) extraction pipelines within
tools/financial_rigor.py. - Structured Sectioning: Uses regex pattern
Item\s+(\d+[A-Z]?)\.\s+([^\n]+)to split documents into standardized headings like "Item 1 – Business". - LLM Integration: Outputs normalized JSON consumed by the prompt template in
skills/earnings-review.md, enforcing source-grounded responses without secondary interpretation.
Frequently Asked Questions
What file types does the earnings-review skill support?
The skill processes both HTML and PDF filings. For SEC EDGAR 10-K and 10-Q reports, it typically parses HTML versions using html_to_text(). For Taiwanese MOPS filings and other jurisdictions that provide PDFs only, it utilizes pdf_to_text() with pdfminer.six to extract content before applying the section-splitting regex.
How does the skill handle different filing formats across countries?
The split_into_sections() function in tools/financial_rigor.py accepts the plain text output regardless of original format (HTML or PDF). While the regex pattern targets SEC-style "Item X" headings, the same logic applies to IO-K and other regional filings by mapping their equivalent heading structures to the same extraction logic, making the skill extensible to new jurisdictions by adjusting the regex or adding jurisdiction-specific heading mappings.
Can I use this parser independently of the LLM skill?
Yes. The core parsing functions—fetch_sec_filing(), html_to_text(), and split_into_sections()—are importable standalone utilities within tools/financial_rigor.py. You can use these to build financial data pipelines, perform bulk extraction of specific sections like "Risk Factors" across multiple years, or integrate the structured output into quantitative analysis workflows without invoking the LLM components.
Where is the prompt template that prevents the LLM from using secondary sources?
The constraint instructing the LLM to use only primary source material resides in skills/earnings-review.md and is duplicated in the auto-generated codex-skills/earnings-review/SKILL.md. These files contain the system prompt that accompanies the structured JSON payload, explicitly forbidding the model from referencing analyst reports, news articles, or internal training data when answering questions about the filing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →