---
title: Bedrock Knowledge Base Cost Starts at the Vector Store
description: A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic. $160 of the smaller bill is Claude Sonnet 5.5.
url: https://www.factualminds.com/blog/bedrock-knowledge-base-cost-optimization-2026/
datePublished: 2026-10-09T00:00:00.000Z
dateModified: 2026-10-09T00:00:00.000Z
author: palaniappan-p
category: Cost Optimization & FinOps
tags: amazon-bedrock, knowledge-bases, cost-optimization, finops, rag
---

# Bedrock Knowledge Base Cost Starts at the Vector Store

> A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic. $160 of the smaller bill is Claude Sonnet 5.5.

On 8 October 2026, AWS Cost Explorer, AWS Budgets, and AWS Cost Management Dashboards added Amazon Bedrock product attributes. You can group and filter Bedrock spend by model, model provider, inference type, and feature, at no extra charge, in every commercial Region. GovCloud and the China Regions are excluded. Attribute history starts on 1 September 2026. A query whose start date is earlier fails with `DataUnavailableException`.

Those attributes show Bedrock lines such as on-demand inference and the reranker. They do not show OpenSearch Serverless or S3 Vectors. On a self-managed knowledge base, the omitted line is often the floor.

The [Knowledge Base cost calculator](/tools/aws-bedrock-knowledge-base-pricing-calculator/) uses us-east-1 rates dated 29 September 2026. One corpus: 5,000 documents, 16 KB average, weekly sync, 50,000 Retrieve calls, 10,000 RetrieveAndGenerate calls on Claude Sonnet 5.5 at 4,000 input tokens and 800 output tokens, rerank off, Data Automation off. S3 Vectors totals **$163.28 a month**. OpenSearch Serverless Classic totals **$510.52**. **$160** of the smaller total is the model. The unrounded gap is $347.23.

## Two prices, two products

The [Bedrock pricing page](https://aws.amazon.com/bedrock/pricing/), checked 9 October 2026, lists two Knowledge Base products.

**Managed Knowledge Bases** charge $5 per GB of raw data per month for the index. Standard Retrieve is $1 per 1,000 API calls. The managed parser, the managed embedding model, and the managed reranker are $0. Agentic retrieval adds $4 per 1,000 Agentic Retrieve calls plus $1 per 1,000 underlying Retrieve calls. If you pick your own embedding or rerank model, that model's price is extra. The site calculator does not model this product.

**Self-managed Knowledge Bases** are the path the calculator does model. You choose the vector store. There is no separate hourly knowledge base fee. You pay embeddings, store, optional parsing, optional rerank, and the generation model.

If the 16 KB figure on this corpus is the raw size, 5,000 documents are about 0.08 GB. Managed index storage at $5 per GB is under $1. The 50,000 Retrieve calls are $50. Generation is still extra if you call a model afterward. Self-managed S3 Vectors, on the calculator's illustrative query rate, prices 60,000 queries at $3. Those two query meters are not the same product. Pick managed when you want AWS to own parsing, embeddings, and the index, and the raw corpus stays small. Pick self-managed S3 Vectors when the documents already live in S3 and Retrieve volume would make $1 per 1,000 calls the line you notice.

## Five meters on a self-managed knowledge base

| Meter | What the calculator uses | This corpus, weekly |
| --- | --- | --- |
| Embeddings | Titan Text Embeddings v2 at $0.02 per 1M tokens, 500 tokens per chunk | $0.11 |
| Parsing | Bedrock Data Automation standard output at $0.010 per page, off by default | $0 |
| Vector store | S3 Vectors illustrative $1/GB-month and $0.05 per 1,000 queries, or OpenSearch Classic | $3.17 or $350.40 |
| Rerank | Amazon Rerank 1.0 at $1 per 1,000 queries, off by default | $0 |
| Generation | Claude Sonnet 5.5 at $2 / $10 per 1M tokens | $160.00 |

The 16 KB documents become 9 chunks each under the calculator's rule (500 tokens, 4 bytes per token), so 5,000 documents are 45,000 vectors and 0.17 GB. Weekly sync re-embeds 25% of those chunks. Daily is 30%. A monthly cycle re-embeds 100%, which is $0.45 of Titan tokens. Moving weekly to daily changes embeddings from $0.11 to $0.14.

The $0.02 per million tokens, the S3 Vectors unit prices, and the $1 rerank rate are the calculator's model as of 29 September 2026. The S3 Vectors storage and query rates are marked illustrative in that file. The $0.010 per page figure is the Bedrock pricing page's own example for Data Automation standard output used as a Knowledge Base parser. Cohere Rerank 3.5, on that same page, is $2 per 1,000 queries, and one query holds at most 100 chunks. A request with 350 chunks is billed as 4 queries.

Sonnet 5.5 at $2 / $10 per million tokens is the rate in the site Bedrock calculator. That module treats it as Anthropic's list price until the Bedrock pricing page prints the Bedrock rate. Re-check the model card before you lock a budget to $160.

OpenSearch compute is 2 OCUs × 730 hours × $0.24 = $350.40. The $0.24 rate is the site OpenSearch calculator. The [OpenSearch pricing page](https://aws.amazon.com/opensearch-service/pricing/), checked the same day, states the Classic minimum in words: at least 2 OCUs for the first collection (1 indexing OCU with primary and standby, 1 search OCU with a replica). The static page text states that minimum and leaves the hourly dollar in the regional price list. A few cents of managed storage on 0.17 GB sit inside the $510.52 total, so the rounded lines ($0.11 + $350.40 + $160.00) land one cent under it. That is also why $510.52 minus $163.28 is $347.24 while the unrounded gap is $347.23.

NextGen collections on that page have no minimum. Indexing and search OCUs scale to zero after 10 minutes of inactivity. A vector search collection still cannot share OCUs with search or time series collections, even on the same KMS key.

## Read the Bedrock lines in Cost Explorer

Group a month of Bedrock by `feature`, then by `model`. The group type in the Cost Explorer API is `PRODUCT_ATTRIBUTE`. The keys documented for Bedrock are `provider`, `model`, `inferenceType`, and `feature`. Examples in the API reference include Claude Sonnet 5, Claude Haiku 4.5, Anthropic, input tokens, output tokens, on-demand inference, and reranker.

`GetCostAndUsage` does not require a service filter for this. Bedrock cost can show up under more than one service name. Omit the service filter so you do not drop a line. When you do filter, the service name has to match exactly, or the call fails validation.

AWS CLI 2 with Cost Explorer support for `PRODUCT_ATTRIBUTE`. The time period has to start on or after 1 September 2026. An older CLI that rejects the group type belongs in the console, which gained the same control on 8 October 2026.

```bash
aws ce get-cost-and-usage \
  --time-period Start=2026-10-01,End=2026-11-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --group-by Type=PRODUCT_ATTRIBUTE,Key=feature
```

Filter one model the way the API reference does. Swap the model string for the name Cost Explorer shows, not the model ID.

```bash
aws ce get-cost-and-usage \
  --time-period Start=2026-10-01,End=2026-11-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --filter '{"ProductAttributes":{"Key":"model","Values":["Claude Sonnet 5"],"MatchOptions":["EQUALS"]}}'
```

Budgets can filter on the same attributes. Build the budget in the console against one model, then attach it to the channel the platform team already reads. Combine the attributes with cost allocation tags on application inference profiles, and with IAM principal tags, when you need the application or the role and not only the model.

Then open the OpenSearch and S3 lines for the same month. If Bedrock attributes sum to roughly the generation total and OpenSearch is sitting near $350, the knowledge base is on a Classic collection.

## Pick the store

Use S3 Vectors for a new self-managed knowledge base. The trade-off is hybrid keyword-plus-vector search, and the latency target already spelled out in the [S3 Vectors post](/blog/amazon-s3-vectors-native-vector-storage/) and the vector store decision. Those jobs stay on OpenSearch or MemoryDB. This page does not restate the June 2026 query limits.

If the job needs OpenSearch, create a NextGen collection and set a maximum OCU. Classic is the type that bills the 2 OCU floor the calculator models at $350.40 a month, including hours with no queries. Vector collections do not share that floor with a search collection on the same key. The first vector collection opens a new pair of OCUs.

Aurora PostgreSQL with pgvector is the right store when the application already runs on Aurora and the vectors should sit next to the rows. The calculator does not price it, so this page has no Aurora total to compare with the $163.28.

Sync the source bucket on change, not on a calendar that re-embeds a stable corpus. The calculator's weekly assumption is 25% of chunks. A change of chunking strategy or embedding model rebuilds the index. On this corpus that rebuild is $0.45 of Titan tokens. Schedule it when the index definition changes.

Turn on Data Automation only for pages with tables, figures, or scans. The pricing page's 1,000-page example at $0.010 is $10. Parsing all 5,000 documents on the weekly 25% assumption adds $12.50. Plain text does not need that parser.

## Code that changes the bill

`RetrieveAndGenerate` always calls the model. `Retrieve` returns chunks and lets you skip generation when the question already has an answer.

boto3 `bedrock-agent-runtime`, Retrieve API. `numberOfResults` caps the chunks. The metadata filter runs against the index before that cap. Five chunks is the starting point for a chat answer.

```python
client = boto3.client("bedrock-agent-runtime", region_name="us-east-1")

retrieved = client.retrieve(
    knowledgeBaseId=kb_id,
    retrievalQuery={"text": question},
    retrievalConfiguration={
        "vectorSearchConfiguration": {
            "numberOfResults": 5,
            "filter": {"equals": {"key": "department", "value": "engineering"}},
        }
    },
)
```

Attach rerank on the second call, after the first pass misses. Pull `rerank_model_arn` from the model card in the console. The field is `modelArn` on `VectorSearchBedrockRerankingModelConfiguration`. Amazon Rerank 1.0 is $1 per 1,000 queries in the calculator. Cohere Rerank 3.5 is $2 per 1,000 queries on the pricing page, with at most 100 chunks in a query. Twenty candidates stay inside one Cohere query. One hundred is the API maximum for `numberOfRerankedResults`.

```python
reranked = client.retrieve(
    knowledgeBaseId=kb_id,
    retrievalQuery={"text": question},
    retrievalConfiguration={
        "vectorSearchConfiguration": {
            "numberOfResults": 20,
            "rerankingConfiguration": {
                "type": "BEDROCK_RERANKING_MODEL",
                "bedrockRerankingConfiguration": {
                    "numberOfRerankedResults": 5,
                    "modelConfiguration": {"modelArn": rerank_model_arn},
                },
            },
        }
    },
)
```

Generate on a smaller model, with the output cap set to the 800 tokens this example already prices. The same 10,000 calls on the calculator's Claude Haiku 4.5 pin ($0.80 / $4.00 per million tokens) are $32 of input and $32 of output, **$64** instead of $160. Keep Sonnet 5.5 for the questions that fail a review set. Nova Lite in the same calculator is $0.06 / $0.24 per million tokens, $4.32 on this token shape, and it is the wrong default for policy text until that review set passes.

boto3 `bedrock-runtime` Converse. `model_id` is the Haiku 4.5 ID from the console. The cache checkpoint sits after the retrieved context and before the question. Default TTL is 5 minutes. Set `ttl` to `1h` only on models that document the one-hour cache. A model that only supports 5 minutes rejects the field.

```python
runtime = boto3.client("bedrock-runtime", region_name="us-east-1")

answer = runtime.converse(
    modelId=model_id,
    system=[{"text": "Answer from the retrieved passages. Cite the passage."}],
    messages=[
        {
            "role": "user",
            "content": [
                {"text": retrieved_passages},
                {"cachePoint": {"type": "default"}},
                {"text": question},
            ],
        }
    ],
    inferenceConfig={"maxTokens": 800},
)
```

On 7 October 2026 the Bedrock pricing page cut Sonnet 5.5 cache-read price by 50% from the prior rate. Confirm the new dollar on the model card before you put a cache-hit percentage in a budget. Caching a context that changes every call does not hit.

## One corpus, two totals

| Path | Embeddings | Vector store | Sonnet 5.5 | Total |
| --- | --- | --- | --- | --- |
| S3 Vectors | $0.11 | $3.17 | $160.00 | $163.28 |
| OpenSearch Classic | $0.11 | $350.40 | $160.00 | $510.52 |

> **What broke** — A console default that accepts the auto-created OpenSearch Serverless collection. Cost Explorer grouped by Bedrock `feature` shows on-demand inference and the reranker, and almost none of the $350.40, because those OCUs bill as OpenSearch. You see it when the Bedrock attribute total and the OpenSearch service total for the same month are different numbers. Recovery is a knowledge base on S3 Vectors, a re-ingest from the source bucket, and deletion of the Classic collection after the new index answers the same questions. Changing the sync from weekly to daily does not close the gap. Embeddings move by about two cents.

> **Reproduce this** — Open the [Bedrock Knowledge Base cost calculator](/tools/aws-bedrock-knowledge-base-pricing-calculator/). Set 5,000 documents, 16 KB average, weekly sync, 50,000 Retrieve calls, 10,000 RetrieveAndGenerate calls, Claude Sonnet 5.5, 4,000 input tokens, 800 output tokens, rerank off, Data Automation off. S3 Vectors totals $163.28 a month. OpenSearch Serverless totals $510.52. The rates are that calculator's us-east-1 model as of 29 September 2026. The Classic 2 OCU minimum matches the OpenSearch pricing page checked 9 October 2026. The $0.24 hourly rate comes from the site OpenSearch calculator. The static pricing-page text states the 2 OCU minimum and leaves the hourly dollar in the regional price list.

Rerank on every one of the 60,000 queries adds $60 and moves the S3 total to $223.28. That is the wrong lever on a corpus whose vector store is already $3.17.

## What to Do This Week

1. In Cost Explorer, group October Bedrock spend by `feature` and by `model`. The period has to start on or after 1 September 2026.
2. List OpenSearch Serverless collections that Bedrock created for knowledge bases. If the collection is Classic and the corpus is small, run it through the calculator before the next sync. Delete the collection only after the replacement index answers the review questions.
3. Call `Retrieve` with `numberOfResults` of 5 and a metadata filter. Turn rerank on for the misses.
4. Move generation to the Haiku 4.5 pin on the same token shape ($64 versus $160) and keep Sonnet 5.5 for the review-set misses.
5. Add a budget filtered to one model. Tag the application inference profile so the attribute view and the tag view name the same application.

## What This Post Doesn't Cover

Creating the knowledge base, chunking choices, and the ingestion clicks live in the [RAG setup guide](/blog/how-to-build-rag-pipeline-amazon-bedrock-knowledge-bases/). S3 Vectors product limits, including the 10,000-result query cap as of 16 June 2026, live in the [S3 Vectors post](/blog/amazon-s3-vectors-native-vector-storage/). AgentCore runtime, memory, and the other eleven components live in the [AgentCore pricing post](/blog/amazon-bedrock-agentcore-pricing-12-components/).

Aurora pgvector dollars are not in the calculator. NextGen OCU behavior under a sustained query rate is not on the pricing page beyond the idle rule (scale to zero after 10 minutes). Source objects in S3 bill at S3 rates and are outside both totals above. GovCloud and the China Regions do not have the product attributes. Managed Knowledge Base agentic retrieval quality is a separate test. This post prices the call. It does not score the answers.

## FAQ

### When should a Knowledge Base stay on OpenSearch Serverless?
Keep OpenSearch when the query needs hybrid keyword plus vector search in one engine, or when the latency target is the one the S3 Vectors post already sends to OpenSearch or MemoryDB. Use a NextGen collection so idle compute can scale to zero after 10 minutes. Classic is the collection type with the 2 OCU minimum the calculator prices at $350.40 a month.

### When should you not turn on rerank or Data Automation?
Leave rerank off when the first Retrieve pass already returns the right chunks. The calculator prices Amazon Rerank 1.0 at $1 per 1,000 queries, so 60,000 queries add $60. Leave Bedrock Data Automation off for plain text. The pricing page charges $0.010 per page for standard output when Data Automation is the Knowledge Base parser. On this corpus, weekly parsing of every document adds $12.50.

### Does Cost Explorer show the whole Knowledge Base bill?
No. Product attributes cover Bedrock lines such as model, provider, inference type, and feature. OpenSearch Serverless and S3 Vectors stay on their own service lines. Compare the Bedrock attribute total with those service totals for the same month.

### Is there a separate hourly fee for a Knowledge Base?
Self-managed Knowledge Bases have no separate hourly fee in the site calculator. You pay embeddings, the vector store, optional parsing and reranking, and model inference. Managed Knowledge Bases are a different price on the Bedrock pricing page: $5 per GB of raw data per month, $1 per 1,000 Retrieve calls, with the managed embedding model and managed reranker included.

### What goes wrong if you group Cost Explorer before September 2026?
Product attribute data is available for time periods that start on or after 1 September 2026. An earlier start date fails with DataUnavailableException. The console control shipped on 8 October 2026.

### When is a full re-ingest the right bill?
Pay it when you change the chunking strategy or the embedding model, because those changes rebuild the index. On this calculator corpus a full monthly re-embed is $0.45 of Titan tokens. Do not schedule that rebuild to save money. The OpenSearch Classic floor is $350.40 either way.

---

*Source: https://www.factualminds.com/blog/bedrock-knowledge-base-cost-optimization-2026/*
