What Is the Patent-OA Sub-Skill? A Complete Guide to Automating Patent Office Action Workflows
The patent-oa sub-skill is a specialized component of the Patent Disclosure skill that automates the entire Patent Office Action (OA) workflow, from ingesting case documents and storing vector embeddings to searching historical cases, generating opinion documents, and redacting sensitive information.
The patent-oa sub-skill lives within the handsomestWei/patent-disclosure-skill repository and provides patent attorneys and IP professionals with a Python-based toolbox for managing Office Action responses. By combining SQLite-based storage, semantic search capabilities, and automated document generation, this sub-skill transforms manual OA preparation into a streamlined, data-driven process.
Core Capabilities of the Patent-OA Sub-Skill
The patent-oa sub-skill addresses five critical stages of the OA response lifecycle:
- Ingestion – Imports raw OA PDFs and playbook documents, converting them into structured markdown while extracting metadata.
- Storage – Persists case information and generates vector embeddings in a local SQLite database for fast semantic retrieval.
- Search – Enables keyword and semantic queries across previously handled cases to find relevant prior art or argument patterns.
- Generation – Converts markdown opinion drafts into professionally formatted DOCX files suitable for filing.
- Redaction – Automatically masks confidential data such as applicant names, claim numbers, and proprietary details before external sharing.
Key Components and Source Files
The sub-skill is organized under skills/patent-oa/tools/ with each module handling a specific aspect of the workflow.
Data Storage and Vector Embeddings
The tools/store.py module initializes a SQLite-based datastore that holds both case metadata and dense vector embeddings. This dual-storage approach allows for fast keyword searches combined with semantic similarity matching.
from patent_oa.tools.store import init_store
# Initialise the OA datastore (creates a SQLite DB if missing)
store = init_store(db_path="~/oa_store.db")
When embedding models change, tools/rebuild_vectors.py recalculates embeddings for all stored cases without requiring re-ingestion of the original documents.
Document Ingestion
Raw Office Actions enter the system through tools/ingest_case.py and tools/ingest_playbook.py. These utilities parse PDF documents, convert content to markdown, extract structural metadata, and optionally generate vector embeddings for semantic search.
from patent_oa.tools.ingest_case import ingest_case
# Ingest a new Office Action PDF
case_id = ingest_case(
pdf_path="D:/OA/2023-07-15_OfficeAction.pdf",
vault_root="D:/Obsidian/MyVault",
store=store,
)
print(f"Case stored with ID: {case_id}")
Case Retrieval
The tools/search_cases.py module provides semantic search capabilities across the case database. It supports queries by examiner name, claim language, dates, or conceptual similarity, returning the top-k most relevant historical cases.
from patent_oa.tools.search_cases import search_cases
# Search for past cases involving a specific examiner
results = search_cases(
query="examiner Smith",
store=store,
top_k=5,
)
for r in results:
print(r["case_id"], r["title"])
Document Generation
Once an opinion is drafted in markdown, tools/emit_opinion_docx.py transforms it into a formatted Microsoft Word document. This bridges the gap between collaborative markdown editing and formal patent office submission requirements.
from patent_oa.tools.emit_opinion_docx import emit_opinion_docx
# Generate a DOCX opinion from markdown template
emit_opinion_docx(
input_md="outputs/oa/案/意见陈述_20260903120000.md",
output_docx="outputs/oa/案/意见陈述_20260903120000.docx",
)
Privacy Protection
Before sharing drafts with clients or external counsel, tools/redact.py processes documents to remove sensitive identifiers. This includes masking applicant names, specific claim numbers, and other confidential metadata that should not circulate beyond the core prosecution team.
from patent_oa.tools.redact import redact_file
# Redact sensitive fields before distribution
redact_file(
input_path="outputs/oa/案/意见陈述_20260903120000.md",
output_path="outputs/oa/案/意见陈述_20260903120000_redacted.md",
)
Configuration Management
The tools/config.py file centralizes environment-specific settings including database paths, embedding model selections (e.g., OpenAI, local transformers), and output directory structures, ensuring consistent behavior across development and production environments.
Integration Workflow
A typical patent-oa workflow proceeds through four stages as implemented in handsomestWei/patent-disclosure-skill:
- Initialize the SQLite store using
init_store()to establish the database schema and vector index. - Ingest new Office Actions via
ingest_case(), which populates the database with searchable content. - Research historical responses using
search_cases()to identify successful argument patterns or relevant prior art citations. - Produce final deliverables by drafting opinions in markdown, generating DOCX files via
emit_opinion_docx(), and sanitizing them withredact_file()before client review.
Summary
- The
patent-oasub-skill automates the complete Patent Office Action workflow within the handsomestWei/patent-disclosure-skill ecosystem. - It stores case data and vector embeddings in SQLite via
store.py, enabling hybrid keyword-semantic search throughsearch_cases.py. - Document ingestion utilities in
ingest_case.pyandingest_playbook.pyconvert PDFs into structured, searchable markdown. - The
emit_opinion_docx.pytool bridges markdown drafting and formal DOCX production required for patent office submissions. - Privacy protection via
redact.pyensures confidential applicant information remains secure during collaborative review.
Frequently Asked Questions
What does patent-oa stand for?
The term patent-oa abbreviates "Patent Office Action," referring to official correspondence from patent examiners that rejects or objects to claims in a patent application. The sub-skill specifically handles the response preparation and management workflow for these documents.
How does the patent-oa sub-skill store case data?
According to the source code in skills/patent-oa/tools/store.py, the system uses a SQLite database to persist case metadata alongside vector embeddings generated from document content. This architecture supports both structured SQL queries and semantic similarity searches without requiring external database servers.
Can the patent-oa sub-skill handle different document formats?
Currently, the ingestion tools in ingest_case.py are optimized for PDF Office Actions, converting them to markdown for processing. The system architecture supports extensibility for additional formats through the modular design of the ingestion pipeline, though the core implementation focuses on PDF-to-markdown workflows.
Is the patent-oa sub-skill suitable for production patent workflows?
Yes, the sub-skill includes production-ready features such as configurable embedding models via config.py, automated redaction for confidentiality, and deterministic case ID generation for audit trails. However, practitioners should validate redaction patterns against their specific jurisdiction's confidentiality requirements before deploying in high-stakes prosecution matters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →