Python Utility Scripts in the patent-disclosure-skill Repository: A Complete Inventory

The patent-disclosure-skill project provides over twenty reusable Python utility scripts organized under scripts/ and skills/*/tools/ directories, offering command-line tools for patent search, document conversion, vault management, and visualization.

The handsomestWei/patent-disclosure-skill repository is an open-source toolkit designed to automate patent disclosure workflows. Its Python utility scripts serve as both standalone CLI programs and importable library modules, enabling developers to search patent databases, convert document formats, and manage Obsidian vaults programmatically.

Repository Layout and Script Organization

The utilities follow a consistent two-tier architecture that separates general-purpose tools from domain-specific helpers.

  • Top-level scripts: Located in scripts/, these are standalone utilities invoked directly from the repository root.
  • Skill-specific tools: Located in skills/<skill-name>/tools/, these modules implement focused features (e.g., patent-search, vault handling, document conversion) and expose a main() entry point for CLI usage.

All utilities adhere to a uniform pattern: they import necessary third-party libraries, define pure functions for core logic, and expose a main(argv: list[str] | None = None) -> int function. This design allows invocation via python -m <module>, direct import, or subprocess calls from higher-level workflows.

Patent Search and Retrieval Utilities

The skills/patent-search/tools/ directory contains the most comprehensive set of CLI tools for interacting with patent databases.

CNIPA Search CLI (skills/patent-search/tools/cnipa_search.py) provides a full-featured command-line client that scrapes the Chinese Patent Office (CNIPA) EP-UB site, parses results, and outputs JSON. It handles authentication, pagination, and result formatting.

python -m skills.patent-search.tools.cnipa_search \
  --type invention \
  --inventor "张三" \
  --title "人工智能" \
  --class "G06F"

The command prints a JSON array of matching EP-UB results and writes a search_report.md file in the default output directory.

Search Configuration Loader (skills/patent-search/tools/search_config.py) reads config.yaml (with fallback defaults) and normalizes pagination parameters for patent-search queries, ensuring consistent API behavior across environments.

Patent Type Normaliser (skills/patent-search/tools/patent_type.py) exposes helpers such as normalize_patent_type() and google_patents_websearch_query(), plus a CLI entry point for quick type conversion.

Derived Query Builder (skills/patent-search/tools/derived_query.py) constructs advanced query strings (e.g., "derived-from" filters) specifically for the CNIPA search UI syntax.

Emit Search Report (skills/patent-search/tools/emit_search_report.py) formats search-result payloads into Markdown reports and saves them with timestamped filenames.

Document Processing and Conversion Tools

Several utilities handle the extraction and transformation of patent documents between formats.

PDF-to-Text Extractor (skills/patent-oa/tools/pdf_text.py) uses pdfminer.six to extract plain text from PDF opinion documents, exposing a main() function for quick conversion.

python -m skills.patent-oa.tools.pdf_text \
  --input path/to/opinion.pdf \
  --output opinion.txt

Docx-to-MD Converter (skills/patent-oa/tools/md_to_docx.py) converts Markdown files into Word .docx format using python-docx logic, preserving formatting for legal documents.

Math-to-OMML Converter (skills/patent-disclosure/tools/math_to_omml.py) transforms LaTeX-style math strings into Office Math Markup Language (OMML), enabling insertion of equations into Word documents for patent applications.

Playbook Builder (skills/patent-oa/tools/playbook.py) reads collections of opinion PDFs, extracts key sections using heuristics, and writes consolidated "playbook" Markdown files for legal analysis.

Vault and Knowledge Management Scripts

For users managing patent disclosures in Obsidian, the repository provides specialized vault management utilities.

Vault Setup for Obsidian (skills/patent-reader/tools/vault/setup_obsidian_vault.py) creates the directory structure and templates required for an Obsidian vault that stores disclosed patents, including standardized folders and configuration files.

Write Obsidian Note (skills/patent-reader/tools/vault/write_patent_obsidian_note.py) generates Markdown notes with front-matter, backlinks, and citations for given patent records.

from skills.patent_reader.tools.vault.write_patent_obsidian_note import write_note

patent_data = {
    "title": "一种基于区块链的身份认证方法",
    "pub_number": "CN112233445A",
    "abstract": "...",
    "link": "https://example.com/patent/112233445",
}
write_note(patent_data, vault_path="~/Obsidian/PatentVault")

This creates a properly formatted markdown file with YAML front-matter, citation links, and structured summaries suitable for knowledge graphs.

Visualization and Diagramming Utilities

The disclosure workflow includes tools for generating visual assets and technical diagrams.

SVG Screenshot Helper (skills/patent-disclosure/tools/svg_screenshot.py) renders SVG elements to PNG using cairosvg, useful for generating presentation assets from vector graphics.

Mermaid Diagram Renderer (skills/patent-disclosure/tools/mermaid_render.py) calls the Mermaid CLI to produce SVG/PNG diagrams from textual flow-chart definitions, automating the creation of process diagrams for patent disclosures.

python -m skills.patent-disclosure.tools.mermaid_render \
  --definition "graph TD; A-->B; B-->C;" \
  --format png \
  --output diagram.png

Patent-Application Figure Composer (skills/patent-application/tools/compose_application_figure.py) generates composite PNGs of invention figures, handling automated layout, scaling, and optional captioning for patent submission packages.

Development and Infrastructure Helpers

Several utilities support the development environment and cross-platform compatibility.

Standard-IO UTF-8 Wrapper (skills/patent-search/tools/stdio_utf8.py and identical copies across patent-reader, patent-oa, patent-docket, patent-disclosure, and patent-application tools) guarantees UTF-8 encoding for stdin/stdout across platforms and provides a wrapper around subprocess.run that forces UTF-8 output, preventing encoding errors on Windows systems.

Generate Star-History (scripts/generate-star-history.py) walks the repository's commit history to produce JSON summaries of GitHub star activity, useful for generating project-status dashboards and metrics reports.

CAD Virtual-Env Bootstrapper (skills/patent-disclosure/tools/bootstrap_cad_venv.py) automates creation of a Python virtual environment pre-loaded with CAD-related packages (e.g., cadquery), supporting mechanical drawing workflows for patent illustrations.

Audit Claims Helper (skills/patent-application/tools/audit_claims.py) parses claim-tree JSON structures, validates internal consistency, and prints concise audit reports to catch claim dependency errors before filing.

Summary

  • The patent-disclosure-skill repository organizes Python utility scripts into scripts/ (top-level) and skills/*/tools/ (domain-specific) directories.
  • Patent search utilities like cnipa_search.py provide CLI access to Chinese patent databases with JSON output and Markdown reporting.
  • Document processors handle PDF extraction (pdf_text.py), Word conversion (md_to_docx.py), and mathematical notation transformation (math_to_omml.py).
  • Obsidian vault tools automate knowledge base setup and note generation for patent portfolios.
  • Visualization scripts generate diagrams and composite figures using Mermaid, CairoSVG, and PIL.
  • All utilities follow a consistent architectural pattern with main() entry points, supporting direct execution, module import, or subprocess invocation.

Frequently Asked Questions

How do I run the CNIPA patent search utility from the command line?

Execute the module directly using Python's -m flag: python -m skills.patent-search.tools.cnipa_search --type invention --inventor "Name". The script handles authentication, executes the search against the Chinese Patent Office EP-UB site, and outputs structured JSON while automatically saving a Markdown report to the default output directory.

What is the purpose of the stdio_utf8.py files found in multiple skill directories?

These files provide cross-platform UTF-8 encoding enforcement for stdin and stdout streams, ensuring that Chinese characters and special symbols render correctly on Windows systems. They also wrap subprocess.run calls to guarantee UTF-8 output decoding, preventing UnicodeDecodeError exceptions when processing patent documents containing international text.

Can the Obsidian vault tools be used independently of the full skill workflow?

Yes. Both setup_obsidian_vault.py and write_patent_obsidian_note.py expose importable functions and CLI interfaces. You can import write_note() directly into your Python scripts to generate patent notes programmatically, or run the setup script standalone to initialize a fresh Obsidian vault with the proper directory structure and templates for patent disclosure management.

Which utility should I use to convert patent opinion PDFs into editable text?

Use skills/patent-oa/tools/pdf_text.py, which leverages pdfminer.six to extract clean plain text from PDF opinion documents. It exposes a main() function suitable for CLI usage with --input and --output arguments, producing UTF-8 encoded text files ready for downstream NLP processing or diff comparison.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →