How Formulas Are Handled in the Patent Disclosure Process: A Technical Deep Dive

Formulas in the patent disclosure process undergo a rigorous four-stage pipeline—chemical parsing, semantic validation, secure evaluation, and paradigm integration—to ensure scientific accuracy and legal compliance in the handsomestWei/patent-disclosure-skill repository.

The handsomestWei/patent-disclosure-skill repository implements a sophisticated architecture for managing how formulas are handled in the patent disclosure process. This open-source skill treats chemical and mathematical expressions as first-class data structures that require specialized parsing, validation, and rendering workflows. The system separates concerns across four dedicated modules that transform raw formula strings into structured disclosure plans suitable for legal documentation.

The Four-Stage Formula Processing Pipeline

Stage 1: Chemical Syntax Parsing and Normalization

The pipeline begins in skills/patent-disclosure/tools/formula_chem.py, where the parse_formula_atoms() function processes raw chemical strings such as "H₂O" or "C6H12O6·2H₂O". This module normalizes delimiters and extracts precise atomic counts, storing them in a Counter structure. The latex_looks_chemical() function validates that the syntax conforms to chemical notation expectations before further processing occurs.

Stage 2: Semantic Unit and Additive Validation

Next, skills/patent-disclosure/tools/formula_units.py performs semantic validation through the check_additive_units() function. This stage verifies that units and additive scores used within formulas are appropriate for the specific patent case, prohibiting decorative or non-compliant units while ensuring correct weighting parameters. The validation ensures that score weights and measurement units align with patent disclosure standards.

Stage 3: Secure Arithmetic Evaluation

The skills/patent-disclosure/tools/formula_eval.py module handles mathematical computation through the eval_equation() function. Rather than using unsafe evaluation methods, this implementation employs a sandboxed eval environment restricted to a whitelisted set of functions including min and max. This approach safely calculates arithmetic expressions embedded within formula plans without exposing the system to code injection vulnerabilities.

Stage 4: Paradigm Integration and Plan Assembly

The final stage involves skills/patent-disclosure/tools/formula_paradigms.py and skills/patent-disclosure/tools/check_formula_plan.py. The load_paradigms() function loads predefined formula paradigms from YAML or JSON configurations, while paradigm_by_id() retrieves specific interpretation schemas. The check_formula_plan.py module orchestrates the assembly of a comprehensive formula_plan that consolidates parsed atoms, validated units, evaluated results, and selected paradigms into a structured object that drives downstream document generation.

Complete Workflow Implementation

The following Python example demonstrates the complete formula processing workflow as implemented in the patent-disclosure-skill repository:

from formula_chem import parse_formula_atoms, latex_looks_chemical
from formula_units import check_additive_units
from formula_eval import eval_equation
from formula_paradigms import load_paradigms, paradigm_by_id

# 1️⃣ Parse the raw chemical formula string

atoms, err = parse_formula_atoms("C6H12O6·2H2O")
assert not err, f"Parse error: {err}"
print(atoms)       # Counter({'C': 6, 'H': 14, 'O': 8})

# 2️⃣ Validate chemical syntax conventions

assert latex_looks_chemical("C6H12O6·2H2O")

# 3️⃣ Check additive units and score weights

units_ok = check_additive_units({"score": 5, "weight": 0.2})
assert units_ok

# 4️⃣ Evaluate arithmetic components safely

value = eval_equation("5 * 0.2 + 3")
print(value)       # 4.0

# 5️⃣ Load paradigms and retrieve case-specific interpretation

paradigms = load_paradigms()
para = paradigm_by_id(paradigms, "weighted_sum")
print(para["description"])

This workflow transforms raw formulas into structured formula_plan objects that downstream tools consume for Markdown-to-Docx conversion and SVG rendering.

Summary

Frequently Asked Questions

What is the specific role of formula_chem.py in processing patent formulas?

The skills/patent-disclosure/tools/formula_chem.py module serves as the entry point for chemical formula processing, providing the parse_formula_atoms() function to normalize delimiters and extract atomic counts. It also includes latex_looks_chemical() to validate that raw strings conform to chemical notation standards before they enter the validation pipeline.

How does the patent-disclosure-skill prevent security risks during formula evaluation?

The system utilizes eval_equation() in skills/patent-disclosure/tools/formula_eval.py, which implements a sandboxed evaluation environment restricted to a whitelisted set of safe functions such as min and max. This approach prevents code injection while allowing necessary arithmetic calculations for formula-derived values.

What purpose do paradigms serve in the formula disclosure workflow?

Paradigms, managed by skills/patent-disclosure/tools/formula_paradigms.py, are predefined YAML/JSON configurations that specify how formulas should be interpreted in particular patent cases. The paradigm_by_id() function retrieves these schemas, enabling check_formula_plan.py to generate context-aware formula_plan objects that guide downstream document generation and rendering.

Which module coordinates the entire formula processing pipeline?

The skills/patent-disclosure/tools/check_formula_plan.py module acts as the orchestration layer, integrating outputs from formula_chem.py, formula_units.py, and formula_eval.py with paradigm definitions to produce comprehensive formula_plan structures. These plans subsequently drive text generation and illustration tools for final patent document assembly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →