How the audit_claims.py Script Works in the Patent-Application Skill

The audit_claims.py script analyzes patent claim structure by parsing input files, verifying hierarchical dependencies between independent and dependent claims, detecting drafting errors like repetitive phrasing or missing "wherein" clauses, and cross-referencing terminology against the specification to generate compliance reports with appropriate exit codes for CI integration.

The audit_claims.py module serves as the core validation engine within the patent-application skill of the handsomestWei/patent-disclosure-skill repository. This Python script automates quality assurance for patent drafts by processing claim text through a seven-stage pipeline that enforces formal patent office standards. Understanding its architecture enables developers to extend auditing capabilities or integrate automated patent validation into existing workflows.

Core Architecture and Execution Flow

The script implements a sequential processing pipeline defined in skills/patent-application/tools/audit_claims.py. Each stage transforms raw claim text into structured audit results while maintaining strict Unicode compatibility.

Data Ingestion and Normalization

The script begins by invoking stdio_utf8 from skills/patent-application/tools/stdio_utf8.py to read claim files with robust UTF-8 encoding support. This prevents character corruption when processing international applications containing technical symbols or non-ASCII characters. The parser splits input into individual claim objects, strips leading and trailing whitespace, normalizes embedded line breaks, and preserves claim numbering to create a clean representation for analysis.

Hierarchy Verification and Pitfall Detection

Following normalization, the script verifies claim hierarchy by ensuring dependent claims correctly reference preceding independent claims. The analyzer confirms that claim numbering follows a monotonic sequence without gaps. Simultaneously, the detection routine scans for repetitive wording across claims, identifies over-broad language likely to trigger examiner objections, and flags missing essential elements such as "wherein" clauses that define specific embodiments.

Specification Cross-Referencing

Using the lexicon helper from skills/patent-application/tools/lexicon.py, the script extracts domain-specific vocabulary from the patent specification. It validates that each claim incorporates at least one technical term from this lexicon, ensuring strong alignment between the claims and the detailed description. This cross-reference check strengthens the patent application's defensibility by maintaining consistent terminology throughout the document.

Helper Modules and Integration Architecture

The script relies on four specialized utilities located in the skills/patent-application/tools/ directory to perform its functions:

  • stdio_utf8.py: Provides cross-platform UTF-8 I/O operations for reading claim files without encoding errors on diverse operating systems.
  • lexicon.py: Supplies domain-specific vocabulary extracted from the invention description to validate term coverage across all claims.
  • iteration_dialog_log.py: Records each audit step to support conversational interfaces, enabling the skill to report specific issues such as "Claim 3 is missing a 'wherein' clause" to voice assistants or chat interfaces.
  • emit_application_docx.py: Formats audit findings into Microsoft Word documents via emit_application_docx, generating attorney-ready reports that document detected violations and suggested corrections.

Execution Flow and Exit Code Handling

The script follows this operational sequence as implemented in the source:


# Conceptual execution flow based on audit_claims.py implementation

claims_list = load_claims(file_path)                 # Uses stdio_utf8

clean_claims = normalize(claims_list)                # Whitespace/formatting

hierarchy_ok = verify_hierarchy(clean_claims)      # Dependency validation

issue_list = detect_issues(clean_claims)           # Drafting error detection

missing_terms = cross_reference(clean_claims, lexicon)  # Terminology check

report_file = generate_report(issue_list, missing_terms)  # DOCX generation

exit(status=0 if not fatal_issues else 1)          # CI/CD signaling

The exit code handling ensures that downstream automation tools can abort patent draft generation when fatal errors occur. A return value of 0 indicates successful validation, while non-zero exits signal critical violations that prevent non-compliant applications from entering filing workflows.

Practical Usage Examples

Command-Line Execution

Run the script directly against a Markdown claim file:

python -m skills.patent-application.tools.audit_claims --input claims.md

This command prints a console summary and writes claims_audit_report.docx via the emit_application_docx utility for attorney review.

CI Pipeline Integration

Incorporate the audit into automated quality gates:

python skills/patent-application/tools/audit_claims.py claims.txt && echo "Claims passed"

The pipeline fails with exit code 1 if the script detects hierarchy violations or missing specification terms, blocking faulty drafts from proceeding to publication stages.

Voice Assistant Integration

When invoked through the skill's conversational interface, iteration_dialog_log captures each audit step, allowing the assistant to verbally summarize critical issues. For example, the system can report that "Claim 2 repeats the phrase 'connected to' three times" without requiring the user to parse technical log files.

Summary

  • The audit_claims.py script operates within handsomestWei/patent-disclosure-skill to validate patent claim structure, terminology, and compliance with patent office standards.
  • File processing relies on stdio_utf8.py for Unicode-safe input handling, supporting international applications with complex technical symbols.
  • Validation includes hierarchy verification, drafting error detection, and cross-referencing against the lexicon.py terminology database.
  • Output formats include console reports and Word documents generated via emit_application_docx.py for legal review.
  • Non-zero exit codes enable CI/CD automation to reject non-compliant patent drafts before filing.

Frequently Asked Questions

What file formats does audit_claims.py support for input claims?

The script primarily processes plain text and Markdown files containing patent claims. The stdio_utf8 utility ensures proper handling of UTF-8 encoded characters, supporting international patent applications with technical symbols and non-ASCII text without corruption.

How does the script distinguish between independent and dependent claims?

The analyzer examines claim references during the hierarchy verification stage. It checks that dependent claims explicitly reference a preceding claim number using phrases such as "The device of claim 1..." and validates that numbering follows a continuous, monotonic sequence without gaps or duplicates.

Can audit_claims.py integrate with continuous integration pipelines?

Yes. The script returns exit code 0 when no fatal errors are detected and exit code 1 when critical violations occur. This behavior allows CI systems to halt builds when patent claims violate structural rules or lack required terminology from the specification lexicon.

What specific drafting errors does the script detect?

The detection routine identifies repetitive wording across multiple claims, flags over-broad language that may trigger patent examiner rejections, and identifies missing essential elements such as "wherein" clauses. These checks ensure each claim precisely defines the invention's scope while maintaining consistency with the detailed specification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →