Does Patent-OA Skill Support Vector Retrieval from a Case Library? Complete Technical Analysis

Yes, the patent-oa skill provides full vector retrieval capabilities for its case library through dense embeddings generated by Sentence-Transformers and cosine similarity search.

The patent-oa skill in the handsomestWei/patent-disclosure-skill repository implements a production-ready semantic search pipeline. Rather than relying on keyword matching, it encodes patent cases into high-dimensional vectors and retrieves the most relevant precedents through mathematical similarity computation. This architecture enables nuanced matching of technical concepts even when terminology differs between queries and stored cases.

How Vector Retrieval Works in Patent-OA

The skill separates concerns across four specialized modules. Each handles a distinct phase of the vector lifecycle: generation, configuration, search execution, and reusable embedding utilities.

Vector Generation Pipeline

The rebuild_vectors.py script processes the entire case library in batch. Located at skills/patent-oa/tools/rebuild_vectors.py, this module:

  • Discovers all JSON case files in config.CASES_DIR using glob
  • Loads the Sentence-Transformer model (default: all-MiniLM-L6-v2)
  • Encodes each case's text field into a 384-dimensional embedding
  • Persists the mapping as JSON to config.VECTORS_PATH
from sentence_transformers import SentenceTransformer

def build_vectors(cases: List[Dict], model_name: str = "all-MiniLM-L6-v2") -> Dict:
    """
    Generate vector embeddings for each case using the specified sentence transformer model.
    """
    model = SentenceTransformer(model_name)
    vectors = {}
    for case in cases:
        case_id = case.get("id")
        text = case.get("text", "")
        embedding = model.encode(text).tolist()
        vectors[case_id] = embedding
    return vectors

The embedding dimensionality and model architecture are fixed by the all-MiniLM-L6-v2 selection, a compact model optimized for semantic similarity tasks with strong performance on technical text.

Vector Search Implementation

The search_cases.py module at skills/patent-oa/tools/search_cases.py executes runtime retrieval. Its search_cases() function:

  • Loads pre-computed vectors from disk (fast cold-start)
  • Encodes the user query using the identical model configuration
  • Computes cosine similarity via sentence_transformers.util.cos_sim
  • Returns the top_k most similar case objects
from sentence_transformers import SentenceTransformer, util

def search_cases(query: str, top_k: int = 5) -> List[Dict]:
    """
    Search the case library using vector similarity and return the top k most relevant cases.
    """
    # Load vectors and cases

    with open(config.VECTORS_PATH, "r", encoding="utf-8") as f:
        vectors = json.load(f)
    
    # Encode query into vector

    model = SentenceTransformer(config.MODEL_NAME)
    query_vec = model.encode([query])
    
    # Compute similarities

    case_ids = list(vectors.keys())
    case_embeddings = [vectors[cid] for cid in case_ids]
    similarities = util.cos_sim(query_vec, case_embeddings)[0]
    
    # Get top k indices

    top_k_idx = similarities.argsort(descending=True)[:top_k]
    results = [cases[case_ids[i]] for i in top_k_idx]
    return results

The use of util.cos_sim ensures numerical stability and GPU acceleration when available. The descending sort guarantees highest-similarity results surface first.

Reusable Embedding Utility

For components needing ad-hoc vectorization without the full search infrastructure, embed.py at skills/patent-oa/tools/embed.py exposes a thin wrapper:

from sentence_transformers import SentenceTransformer
from . import config

def embed_texts(texts: List[str]) -> List[List[float]]:
    """
    Encode a list of texts into vector embeddings using the configured model.
    """
    model = SentenceTransformer(config.MODEL_NAME)
    embeddings = model.encode(texts).tolist()
    return embeddings

This enables opinion generation modules, similarity thresholding, or cross-reference matching to operate in the same vector space as the retrieval system.

Centralized Configuration

The config.py file at skills/patent-oa/tools/config.py unifies paths and model selection:

  • CASES_DIR: Source directory for JSON case files
  • VECTORS_PATH: Destination for the generated embedding cache
  • MODEL_NAME: Transformer model identifier (default all-MiniLM-L6-v2)

This design prevents drift between vector generation and search—both phases reference identical configuration constants.

Practical Usage Examples

Rebuilding the Vector Index

Execute after adding or modifying cases in the library:

from skills.patent_oa.tools.rebuild_vectors import main

if __name__ == "__main__":
    main()  # Reads CASES_DIR → builds embeddings → writes VECTORS_PATH

The process is idempotent: re-running overwrites VECTORS_PATH with fresh embeddings using the current model configuration.

Performing Semantic Case Retrieval

from skills.patent_oa.tools.search_cases import search_cases

query = "method for reducing power consumption in micro-LED displays"
results = search_cases(query, top_k=3)

for case in results:
    print(f"Case ID: {case['id']}")
    print(f"Title: {case.get('title')}")
    print(f"Similarity: contextual match via vector proximity\n")

Unlike keyword search, this retrieves cases discussing "energy-efficient illumination systems" or "low-power emissive displays" even without exact term matches.

Embedding Arbitrary Patent Text

from skills.patent_oa.tools.embed import embed_texts

claims = [
    "A neural network accelerator with sparse matrix support",
    "An optical waveguide coupling structure for photonic chips"
]
vectors = embed_texts(claims)

# Returns: List[List[float]] with 384-dimensional vectors

Key Source Files and Responsibilities

File Path Primary Function Critical Dependencies
skills/patent-oa/tools/rebuild_vectors.py Batch embedding generation for case library sentence_transformers.SentenceTransformer, glob, json
skills/patent-oa/tools/search_cases.py Runtime vector similarity search sentence_transformers.util.cos_sim, config module
skills/patent-oa/tools/embed.py Ad-hoc text vectorization utility sentence_transformers.SentenceTransformer
skills/patent-oa/tools/config.py Path and model configuration os for environment-aware defaults

Technical Characteristics

  • Embedding model: all-MiniLM-L6-v2 (384 dimensions, ~80MB)
  • Similarity metric: Cosine similarity via optimized tensor operations
  • Storage format: JSON-serialized Python dictionaries
  • Query latency: Sub-second for typical case libraries (disk I/O bound)
  • Update strategy: Full rebuild required; no incremental embedding updates implemented

Summary

  • Vector retrieval is fully supported: The patent-oa skill implements complete dense retrieval from case libraries through four coordinated modules
  • Sentence-Transformers powers embeddings: The all-MiniLM-L6-v2 model generates 384-dimensional semantic representations
  • Cosine similarity drives ranking: util.cos_sim computes query-to-case relevance with GPU acceleration support
  • Configuration is centralized: config.py ensures generation and search use identical model and path settings
  • Utilities are modular: embed.py enables vector operations beyond the core search pipeline

Frequently Asked Questions

What embedding model does patent-oa use for vector retrieval?

The skill defaults to all-MiniLM-L6-v2, a 22.7M parameter Sentence-Transformer model. This produces 384-dimensional embeddings optimized for semantic similarity tasks. The model identifier is stored in config.MODEL_NAME and referenced by rebuild_vectors.py, search_cases.py, and embed.py to ensure consistent vector spaces across all operations.

Is vector search in patent-oa performed in real-time or pre-computed?

Both phases are separated for efficiency. Case embeddings are pre-computed by rebuild_vectors.py and serialized to JSON. At query time, search_cases.py loads these cached vectors and performs real-time encoding of only the user query, then computes similarity against the stored embeddings. This hybrid approach minimizes latency while maintaining semantic accuracy.

Can patent-oa search handle incremental case additions without full rebuilds?

The current implementation in rebuild_vectors.py performs full batch regeneration of all vectors. There is no incremental update logic—new cases require running main() to rebuild the complete VECTORS_PATH file. For large libraries, this suggests an opportunity for optimization through append-only vector stores like FAISS or vector databases.

What similarity threshold does patent-oa use for case retrieval?

The source code does not apply an explicit similarity threshold. The search_cases() function returns exactly top_k results ranked by cosine similarity descending, regardless of absolute score. Downstream consumers must implement threshold filtering if low-relevance matches should be discarded.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →