How to Search for Similar Cases and Draft Responses for Office Actions in Patent Disclosure Skill

Use the search_cases.py module for tag-based or vector semantic search of prior cases, then run ingest_case.py with the --pdf flag to generate a markdown draft response enriched with matching case references.

The Patent Disclosure Skill in the handsomestWei/patent-disclosure-skill repository provides a command-line toolkit for patent professionals to streamline Office Action workflows. According to the source code, the system combines PDF text extraction, configurable vector embeddings, and SQLite-backed case storage to deliver both traditional tag filtering and modern semantic search capabilities.

Understanding the Core Workflow Architecture

The patent workflow follows a clear data pipeline implemented across five primary modules under tools/oa/:

  1. Configuration (config.py) – manages YAML-based settings and API presets
  2. Embedding (embed.py) – abstracts multiple embedding providers for semantic search
  3. Storage (store.py) – SQLite persistence with optional vector columns
  4. Search (search_cases.py) – orchestrates query resolution and retrieval
  5. Drafting (ingest_case.py) – assembles final markdown responses

The entry points expose this functionality through Python module execution:

Capability Command
Search similar cases python -m tools.oa.search_cases …
Generate draft response python -m tools.oa.ingest_case …
Configure settings python -m tools.oa.config …

Configuring Search Capabilities

Enabling Vector Search for Semantic Similarity

Before running semantic searches, configure your embedding provider. The config.py module stores settings in ~/Documents/patent-disclosure-skill/oa/embedding.config.yaml and supports multiple provider presets:

python -m tools.oa config set \
    --preset zhipu \
    --api-key YOUR_ZHIPU_API_KEY \
    --run-selftest

The --run-selftest flag triggers the run_selftest() function, which embedding a short test query and reports result size and timing to verify connectivity. Supported presets in the PRESETS dictionary include Zhipu, DashScope, OpenAI, MiniMax, and local (Sentence-Transformers).

Disabling Vectors for Tag-Only Operation

When API quotas are limited or operating offline, force tag-only mode:

python -m tools.oa config skip-vector

This sets vector_enabled: false in the configuration YAML.

Searching for Similar Cases

The search_cases.py module provides flexible input handling through its resolve_query() function. Users can supply:

  • --pdf <file> – automatically extract text via pdf_text.read_document()
  • --query <text> – raw query string
  • --query-file <path> – file containing query text
  • Tag filters: --tag, --statute, --defect, --domain, --patent-type

Tag-Only Search Example

For traditional metadata filtering without embeddings:

python -m tools.oa.search_cases \
    --tag "G06F" \
    --defect "lack_of_inventive_step" \
    --patent-type "utility" \
    --top-k 5

Output includes JSON with retrieval_mode: "tags" and matching cases with diff_fields showing field-by-field alignment.

Semantic Search with PDF Input

To search by semantic similarity to an Office Action:

python -m tools.oa.search_cases \
    --pdf ./office_action.pdf \
    --statute "35 U.S.C. § 103" \
    --top-k 10

The execution flow in search_cases.py handles this as follows:

  1. resolve_query() extracts and truncates text to _DEFAULT_QUERY_CHARS
  2. If vector mode enabled, Embedder.embed_one() generates query embedding
  3. On EmbedError (timeout, missing key, network failure), falls back to tags-only
  4. store.search() executes with vector (or None) plus filters
  5. Returns JSON with query_preview, retrieval_mode, vector_enabled, embed_error if applicable, and hits array

The embed.py module's EmbedError exception class enables graceful degradation when embedding services fail.

Drafting Responses to Office Actions

The ingest_case.py module's build_draft_from_pdfs() function combines PDF extraction, case search, and markdown templating into a complete draft workflow.

Generating a Draft with Auto-Referenced Cases

python -m tools.oa.ingest_case \
    --pdf ./office_action.pdf \
    --draft-only \
    --top-k 8

This command executes the following pipeline as implemented in the source:

  1. Extracts raw text via pdf_text.read_document()
  2. Optionally runs similarity search (reusing search_cases.py logic)
  3. Merges extracted text with markdown template
  4. Inserts References section populated from matching cases
  5. Writes output to oa/drafts/<case_id>.md in your Obsidian vault

The JSON response contains "draft_md" with the absolute path and embedded case references ready for attorney review and editing.

Database and Embedding Internals

SQLite Case Store (store.py)

The store.py module manages oa_vectors.sqlite with two primary functions:

  • open_store(mode) – opens connection, optionally requiring vector column when mode == "vector"
  • search(vector, **filters) – returns top-K matches combining vector similarity with tag/statute/defect/domain/patent-type filters

Vector columns store pre-computed embeddings for the case corpus, enabling sub-second similarity retrieval.

Multi-Provider Embedding (embed.py)

The Embedder class in embed.py unifies three implementation strategies:

Provider Type Implementation Methods
local Sentence-Transformers embed_one(), embed_texts()
openai_compatible Zhipu/DashScope/OpenAI Same interface, different endpoints
minimax MiniMax API Custom normalization

Batch processing via embed_texts() improves throughput for indexing operations, while embed_one() serves interactive queries.

Summary

  • Configure first – run python -m tools.oa config set with your API provider to enable vector search, or skip-vector for offline operation
  • Search flexibly – use search_cases.py with --pdf for semantic similarity, or tag filters for precise metadata matching
  • Draft efficiently – ingest_case.py with --draft-only generates Obsidian-ready markdown with auto-inserted case references
  • Handle failures gracefully – the system automatically falls back to tag-only search when embedding services fail
  • Customize storage – SQLite-backed store.py supports both vector and non-vector modes via open_store(mode)

Frequently Asked Questions

What file formats are supported for Office Action input?

The pdf_text.py module handles PDF extraction exclusively. Both search_cases.py and ingest_case.py accept PDF files via the --pdf flag, which calls pdf_text.read_document() internally. For text files or direct strings, use --query-file or --query respectively.

How does the system handle embedding API failures?

The embed.py module raises EmbedError for timeouts, missing API keys, and network problems. The search_cases.py orchestrator catches these exceptions, records the error in the embed_error output field, and automatically falls back to tag-only retrieval mode with retrieval_mode: "tags".

Can I use local embeddings without cloud API costs?

Yes. The Embedder class in embed.py supports a local preset using Sentence-Transformers. Configure with python -m tools.oa config set --preset local to run entirely offline, though initial model download and slower inference apply compared to cloud providers.

Where are draft responses saved and how are they organized?

The build_draft_from_pdfs() function in ingest_case.py writes markdown files to oa/drafts/<case_id>.md within your Documents folder. The filename derives from the case identifier extracted from the PDF or provided metadata, and the directory structure integrates with standard Obsidian vault layouts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →