# How to Generate kNN Vector Embeddings for Vector Search in Amazon OpenSearch

> Learn to generate kNN vector embeddings in Amazon OpenSearch using Cohere models. Enable semantic similarity search with k-NN queries and hybrid filtering for efficient vector search.

- Repository: [AWS Samples/sample-for-amazon-opensearch-service-tutorials-101](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101)
- Tags: deep-dive
- Published: 2026-02-25

---

**kNN vector embeddings are generated using Amazon Bedrock's Cohere embedding model and stored in OpenSearch's native `knn_vector` field, enabling semantic similarity search through the k-NN query clause with optional hybrid filtering.**

The `aws-samples/sample-for-amazon-opensearch-service-tutorials-101` repository implements an end-to-end vector search pipeline that transforms product metadata into dense vector representations. This solution demonstrates how to generate 1024-dimensional embeddings using fully managed ML models and leverage OpenSearch's approximate nearest neighbor (ANN) engine for low-latency semantic retrieval across both persistent and in-memory storage tiers.

## Generating kNN Vector Embeddings with Amazon Bedrock

### The get_embedding Helper Function

In [`artifacts/index_lambda/opensearch_index.py`](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py), the `get_embedding` function handles all embedding generation by invoking the **Cohere `embed-english-v3`** model through Amazon Bedrock. This function accepts a text string and returns a 1024-dimensional float array representing the semantic meaning of the input.

```python
def get_embedding(text):
    """Gets embedding for text using Cohere model via Bedrock."""
    body = json.dumps({
        "texts": [text],
        "input_type": "search_document",
        "truncate": "END",
        "embedding_types": ["float"]
    })
    response = bedrock_client.invoke_model(
        modelId=MODEL_ID,
        accept='application/json',
        contentType='application/json',
        body=body
    )
    response_body = json.loads(response.get('body').read())
    return response_body['embeddings']['float'][0]

```

*Source:* [opensearch_index.py – get_embedding](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py#L81-L102)

The function sets `input_type` to `"search_document"` when indexing content and implicitly uses query-time embeddings for search (handled by the same function with equivalent payload structure). The returned vector captures semantic relationships between product attributes, enabling similarity matching beyond simple keyword overlap.

## Indexing Embeddings in OpenSearch

### Creating Vector Index Mappings

Before ingesting data, the pipeline establishes dedicated indices with `knn_vector` field mappings. The repository provides two storage configurations:

- **On-disk index** (`create_vector_index_on_disk_mode`): Persists vectors to disk using HNSW (Hierarchical Navigable Small World) algorithm with FAISS engine, suitable for large-scale collections.
- **In-memory index** (`create_vector_index_in_memory_mode`): Stores vectors in RAM for ultra-low latency retrieval on smaller datasets.

Both indices define the `vector_embedding` field with dimension 1024 and cosine similarity space:

```python

# On-disk index mapping structure

"properties": {
    "vector_embedding": {
        "type": "knn_vector",
        "dimension": 1024,
        "method": {
            "name": "hnsw",
            "space_type": "cosinesimil",
            "engine": "faiss"
        }
    }
}

```

*Sources:* [create_vector_index_on_disk_mode](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py#L52-L80), [create_vector_index_in_memory_mode](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py#L18-L38)

### Bulk Indexing Products

The `vectorize_and_index_products` function orchestrates the embedding pipeline. It reads the product catalog, concatenates relevant text fields (title, category, description), generates embeddings via `get_embedding`, and bulk-indexes documents into both storage tiers simultaneously.

```python
def vectorize_and_index_products(event):
    # … read products from source …

    for product in batch:
        combined_text = f"{product['title']}, Category: {product['category']}, Description: {product['description']}"
        vector_embedding = get_embedding(combined_text)
        product['vector_embedding'] = vector_embedding

        # Prepare bulk payload for on-disk index

        bulk_data_on_disk.append({"index": {"_index": VECTOR_INDEX_NAME_ON_DISK, "_id": uuid.uuid4().hex}})
        bulk_data_on_disk.append(product)

        # Prepare bulk payload for in-memory index

        bulk_data_in_memory.append({"index": {"_index": VECTOR_INDEX_NAME_IN_MEMORY, "_id": uuid.uuid4().hex}})
        bulk_data_in_memory.append(product)
    # … execute bulk API call …

```

*Source:* [vectorize_and_index_products](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py#L407-L460)

This dual-indexing strategy allows runtime selection between cost-effective persistent storage and high-performance memory-based search without reprocessing the source data.

## Executing kNN Vector Search Queries

### Pure Vector Search Implementation

The search Lambda in [`artifacts/search_lambda/opensearch_search.py`](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/search_lambda/opensearch_search.py) handles incoming requests via API Gateway. For vector search requests (`"type": "vector_search"`), it generates a query-time embedding and constructs an OpenSearch `knn` query clause targeting the `vector_embedding` field.

```python
elif body["type"] == "vector_search":
    search_text = body["attribute_value"]
    vector_embedding = get_embedding(search_text)
    search_body = {
        "size": 100,
        "_source": {"excludes": ["vector_embedding"]},
        "query": {
            "knn": {
                "vector_embedding": {"vector": vector_embedding, "k": 100}
            }
        }
    }

```

*Source:* [vector_search handling](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/search_lambda/opensearch_search.py#L81-L97)

The query retrieves the top *k*=100 most similar vectors using the same cosine similarity metric defined in the index mapping. The `_source` exclusion prevents returning the large vector arrays to the client, reducing payload size.

### Selecting Storage Modes at Query Time

The Lambda inspects the `"mode"` parameter to route queries to the appropriate index:

```python
if body["mode"] == "on_disk":
    response = ops_client.search(index=VECTOR_INDEX_NAME_ON_DISK, body=search_body)
elif body["mode"] == "in_memory":
    response = ops_client.search(index=VECTOR_INDEX_NAME_IN_MEMORY, body=search_body)

```

*Source:* [search mode dispatch](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/search_lambda/opensearch_search.py#L69-L74)

## Implementing Hybrid Search

For complex retrieval scenarios, the repository supports **hybrid search** that combines kNN vector similarity with traditional Boolean filters. This approach matches semantic intent while respecting hard constraints like category or color attributes.

```python
search_body = {
    "size": 100,
    "_source": {"excludes": "vector_embedding"},
    "query": {
        "hybrid": {
            "queries": [
                {"bool": {"should": should_match_conditions, "minimum_should_match": 1}},
                {"knn": {"vector_embedding": {"vector": vector_embedding, "k": 100}}}
            ]
        }
    },
    "post_filter": {"bool": {"must": should_match_conditions}},
    "search_pipeline": SEARCH_PIPELINE_NAME
}

```

*Source:* [hybrid search construction](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/search_lambda/opensearch_search.py#L32-L55)

The `hybrid` query structure runs both the keyword match and vector search in parallel, then normalizes scores through the specified search pipeline. The `post_filter` ensures results meet all metadata constraints while preserving the semantic ranking.

## Practical Implementation Examples

### Generating an Embedding for Product Data

```python
from artifacts.index_lambda.opensearch_index import get_embedding

product_description = "Women's lightweight running shoe with breathable mesh upper"
embedding = get_embedding(product_description)
print(f"Vector dimension: {len(embedding)}")  # Output: 1024

```

### Indexing a Product Catalog

Invoke the indexing Lambda to process `products_content.jsonl` and populate both vector indices:

```bash
python -m artifacts.index_lambda.opensearch_index vectorize_and_index_products

```

This executes `vectorize_and_index_products`, which automatically chunks the dataset, generates embeddings via Bedrock, and bulk-uploads to OpenSearch.

### Performing Pure kNN Search via API

```json
POST /search
{
  "type": "vector_search",
  "attribute_value": "red waterproof hiking boots",
  "mode": "on_disk"
}

```

The search Lambda converts the query text to a vector embedding and executes the `knn` query against the on-disk index, returning the 100 most semantically similar products.

### Executing Hybrid Vector + Keyword Search

```json
POST /search
{
  "type": "hybrid_search",
  "attribute_value": "comfortable office chair",
  "mode": "in_memory"
}

```

This triggers the hybrid query pipeline, combining vector similarity with metadata filters to return relevant results that match both semantic intent and categorical constraints.

## Summary

- **Embedding Generation**: The `get_embedding` function in [`artifacts/index_lambda/opensearch_index.py`](https://github.com/aws-samples/sample-for-amazon-opensearch-service-tutorials-101/blob/main/artifacts/index_lambda/opensearch_index.py) generates 1024-dimensional vectors using Amazon Bedrock's Cohere `embed-english-v3` model.
- **Index Architecture**: The solution maintains parallel indices for on-disk persistence and in-memory speed, both utilizing OpenSearch's `knn_vector` field type with HNSW/FAISS indexing.
- **Vector Search**: Query-time embeddings feed into native OpenSearch `knn` queries against the `vector_embedding` field, retrieving the top 100 matches by cosine similarity.
- **Hybrid Capability**: The search Lambda supports combining vector similarity with traditional term filters using OpenSearch's hybrid query syntax and search pipelines.

## Frequently Asked Questions

### What embedding model does the repository use for generating kNN vector embeddings?

The implementation uses the **Cohere `embed-english-v3`** model hosted on Amazon Bedrock. This model outputs 1024-dimensional float vectors optimized for semantic similarity tasks, with specific `input_type` parameters distinguishing between document indexing and query generation.

### What is the difference between on-disk and in-memory vector indices in OpenSearch?

The **on-disk index** persists vectors to storage using the FAISS engine with HNSW graphs, enabling cost-effective search across large datasets (millions of vectors). The **in-memory index** loads vectors entirely into RAM, eliminating disk I/O latency for sub-millisecond retrieval on smaller, high-throughput workloads. Both use identical 1024-dimensional mappings but trade cost against latency.

### How does the hybrid search combine keyword and vector similarity?

Hybrid search constructs a query containing both a `knn` clause for vector similarity and a `bool` clause for metadata filtering. OpenSearch's hybrid query processor executes both sub-queries simultaneously, applies normalization through a search pipeline, and blends the relevance scores. The `post_filter` ensures hard constraints (like exact color matches) are preserved in the final result set.

### What is the dimension size of the vector embeddings used in this implementation?

The vector embeddings are **1024 dimensions** wide. This matches the output specification of the Cohere `embed-english-v3` model and is explicitly configured in the OpenSearch index mappings via the `dimension: 1024` parameter in both `create_vector_index_on_disk_mode` and `create_vector_index_in_memory_mode` functions.