# How to Set Up and Configure the Document Knowledge Base for Querying Processed Content

> Learn how to set up and configure the document knowledge base for querying processed content on AWS. Deploy the Bedrock Knowledge Base stack to index your documents.

- Repository: [aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws)
- Tags: how-to-guide
- Published: 2026-02-25

---

**Deploy the Bedrock Knowledge Base stack from [`nested/bedrockkb/template.yaml`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/bedrockkb/template.yaml) with your chosen vector store type, and the automated custom resources will provision the index, create the knowledge base, and start the ingestion job to index your processed documents.**

The AWS Accelerated Intelligent Document Processing (IDP) solution automates the creation of a searchable **document knowledge base** using Amazon Bedrock. Once configured, the system indexes processed documents from your IDP pipeline and exposes them to natural-language queries via GraphQL or the IDP SDK. This guide walks through the complete setup and configuration based on the actual CloudFormation templates and Lambda handlers in the `aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws` repository.

## Step 1: Deploy the Knowledge Base Stack

The accelerator provisions the knowledge base through a CloudFormation/SAM template located at [`nested/bedrockkb/template.yaml`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/bedrockkb/template.yaml). You must specify the vector store type and embedding configuration during deployment.

First, create a parameters file defining your configuration:

```yaml
Parameters:
  LogLevel: INFO
  pVectorStoreType: S3_VECTORS                # or OPENSEARCH_SERVERLESS

  pS3VectorBucketName: my-idp-kb-bucket       # optional – auto‑generated if empty

  pS3VectorIndexName: bedrock-kb-index
  pEmbedModel: amazon.titan-embed-text-v2:0
  pChunkingStrategy: Fixed-size chunking
  pMaxTokens: 300
  pOverlapPercentage: 15
  pKnowledgeBaseBucketName: my-docs-bucket     # bucket containing processed PDFs

  pInputDocumentUploadFolderPrefix: processed/ # prefix where IDP writes documents

```

Deploy using SAM (while the repository's [`publish.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/publish.py) can build nested stacks automatically, direct SAM invocation gives you explicit control):

```bash
sam deploy \
  --stack-name my-idp-kb \
  --template-file nested/bedrockkb/template.yaml \
  --parameter-overrides $(cat params.yaml | grep -v '^#' | tr '\n' ' ') \
  --capabilities CAPABILITY_NAMED_IAM

```

Key configuration options in [`template.yaml`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/template.yaml) include:

- **`pVectorStoreType`** – Choose `OPENSEARCH_SERVERLESS` (default) or `S3_VECTORS` to determine the backend vector store.
- **Conditions** – The template uses `UseS3Vectors` and `UseOpenSearchServerless` (lines 74‑78) to conditionally create resources based on your selection.
- **Custom Resources** – The template defines `S3VectorManagerFunction` and `StartIngestionJobFunction` (lines 73‑150) to handle provisioning and data ingestion automatically.

## Step 2: Configure the Vector Store and Knowledge Base Resources

When you select `S3_VECTORS`, the stack creates the **S3 Vectors manager Lambda** (`S3VectorManagerFunction`) defined in [`nested/bedrockkb/src/s3_vectors_manager/handler.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/bedrockkb/src/s3_vectors_manager/handler.py). This function automates the creation of the underlying storage and the Bedrock knowledge base itself.

The handler performs three critical operations:

1. **Bucket and index creation** – The `create_s3_vector_resources` function calls `create_vector_index` to initialize the S3-based vector store (lines 98‑124).
2. **Knowledge base provisioning** – The `create_knowledge_base_s3_vectors` function (lines 262‑292) constructs the configuration payload and invokes the Bedrock `create_knowledge_base` API.
3. **Output export** – The Lambda logs the generated **KnowledgeBaseId**, which CloudFormation exports as the `KB_ID` environment variable for downstream resources.

For OpenSearch Serverless deployments, the stack provisions the vector store through native CloudFormation resources rather than the custom Lambda, though the knowledge base creation flow remains similar.

## Step 3: Start the Ingestion Job

After the knowledge base and data source are ready, the **Start Ingestion Job custom resource** automatically indexes your documents. This is handled by the Lambda in [`nested/bedrockkb/src/start_ingestion_job_custom_resource/handler.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/bedrockkb/src/start_ingestion_job_custom_resource/handler.py).

The function invokes the Bedrock `start_ingestion_job` API:

```python
def start_ingestion_job(knowledgeBaseId, dataSourceId):
    try:
        CLIENT.start_ingestion_job(
            knowledgeBaseId=knowledgeBaseId,
            dataSourceId=dataSourceId,
            description="Autostart by CloudFormation"
        )
    except Exception as e:
        logger.warning(f"WARN: start_ingestion_job failed.. {e}")

```

This process scans the S3 bucket specified in `pKnowledgeBaseBucketName` (under the `processed/` prefix) and indexes each document into the knowledge base. The function triggers on **Create** and **Update** stack events (lines 43‑50), ensuring your knowledge base stays synchronized with new processed content.

## Step 4: Query the Knowledge Base

Once ingestion completes, you can query the document knowledge base through two interfaces: the AppSync GraphQL resolver or the IDP Python SDK.

### Query via the AppSync Resolver

The resolver Lambda ([`nested/appsync/src/lambda/query_knowledgebase_resolver/index.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/appsync/src/lambda/query_knowledgebase_resolver/index.py)) handles GraphQL queries by calling the Bedrock `retrieve_and_generate` API. It reads the knowledge base ID from the `KB_ID` environment variable (populated from CloudFormation outputs).

```python
input = {
    "input": {"text": query},
    "retrieveAndGenerateConfiguration": {
        "knowledgeBaseConfiguration": {
            "knowledgeBaseId": KB_ID,
            "modelArn": MODEL_ARN,
        },
        "type": "KNOWLEDGE_BASE"
    }
}
resp = KB_CLIENT.retrieve_and_generate(**input)

```

You can optionally configure guardrails by setting the `GUARDRAIL_ID_AND_VERSION` environment variable. The implementation details are found in lines 35‑78 of [`query_knowledgebase_resolver/index.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/query_knowledgebase_resolver/index.py).

### Query via the IDP Python SDK

For programmatic access, use the `SearchOperation` class from the IDP SDK:

```python
from idp_sdk import IDPClient

client = IDPClient(stack_name="my-idp-stack", region="us-east-1")
search = client.search()
result = search.query(
    question="What is the loan amount for document 123?",
    document_ids=None,          # optional filter

    limit=5
)

print(result.answer)
for c in result.citations:
    print(f"- {c.document.document_id}: {c.text}")

```

The SDK's `SearchOperation.query()` method (defined in [`lib/idp_sdk/idp_sdk/operations/search.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/lib/idp_sdk/idp_sdk/operations/search.py), lines 18‑46) internally instantiates a `SearchProcessor` (from [`lib/idp_sdk/idp_sdk/core/search_processor.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/lib/idp_sdk/idp_sdk/core/search_processor.py)) that formats the request for Bedrock and returns a typed `SearchResult` object.

## End-to-End Deployment Example

Combine all steps to deploy and query your knowledge base:

```bash

# Deploy the stack

sam deploy \
  --stack-name my-idp-kb \
  --template-file nested/bedrockkb/template.yaml \
  --parameter-overrides $(cat params.yaml | grep -v '^#' | tr '\n' ' ') \
  --capabilities CAPABILITY_NAMED_IAM

# Verify ingestion job completion in CloudWatch logs for StartIngestionJobFunction

```

Then query using Python:

```python
from idp_sdk import IDPClient

client = IDPClient(stack_name="my-idp-stack", region="us-east-1")
result = client.search().query("What is the total amount due?")
print(result.answer)

```

## Summary

- **Deployment**: Use [`nested/bedrockkb/template.yaml`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/nested/bedrockkb/template.yaml) with parameters like `pVectorStoreType` and `pEmbedModel` to define your knowledge base architecture.
- **Automation**: Custom-resource Lambdas in [`s3_vectors_manager/handler.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/s3_vectors_manager/handler.py) and [`start_ingestion_job_custom_resource/handler.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/start_ingestion_job_custom_resource/handler.py) automatically provision storage and index documents without manual intervention.
- **Configuration**: The system supports both **OpenSearch Serverless** and **S3 Vectors** backends, selected via the `pVectorStoreType` parameter.
- **Querying**: Access indexed content through the **AppSync resolver** (GraphQL) or the **IDP SDK's** `SearchOperation.query()` method, both leveraging the `KB_ID` environment variable to target the correct Bedrock knowledge base.

## Frequently Asked Questions

### What vector store options are available for the document knowledge base?

The accelerator supports two vector store types configured via the `pVectorStoreType` parameter: **OpenSearch Serverless** (default) or **S3 Vectors**. OpenSearch Serverless provisions a managed vector engine, while S3 Vectors stores embeddings in a dedicated S3 bucket with a custom index managed by the `s3_vectors_manager` Lambda.

### How does the ingestion job know when to start?

The `StartIngestionJobFunction` Lambda triggers automatically on CloudFormation **Create** and **Update** events after the knowledge base and data source resources are ready. It calls the Bedrock `start_ingestion_job` API to begin scanning the document bucket, ensuring processed files are indexed immediately after stack deployment.

### Where is the Knowledge Base ID stored for query operations?

CloudFormation exports the **KnowledgeBaseId** as an environment variable named `KB_ID` to the AppSync resolver Lambda ([`query_knowledgebase_resolver/index.py`](https://github.com/aws-solutions-library-samples/accelerated-intelligent-document-processing-on-aws/blob/main/query_knowledgebase_resolver/index.py)) and makes it available to the IDP SDK. The resolver reads this variable at runtime to construct Bedrock `retrieve_and_generate` requests, while the SDK retrieves it from the stack outputs during client initialization.

### Can I filter queries to specific documents?

Yes. When using the IDP SDK, pass a list of `document_ids` to the `SearchOperation.query()` method to restrict the search scope. The SDK passes these filters to the underlying Bedrock knowledge base, allowing you to query specific subsets of your processed content rather than the entire corpus.