---
title: Bedrock Knowledge Base Cost for Gmail and Google Drive
description: Checked 10 October 2026: managed Bedrock knowledge base storage is $5 per GB of raw data a month. One thousand GB is $5,000 before retrieval. An e-commerce workspace of terabytes of Drive files and millions of Gmail messages reached hundreds of dollars a day after ingestion started.
url: https://www.factualminds.com/blog/bedrock-knowledge-base-cost-gmail-google-drive/
datePublished: 2026-10-10T00:00:00.000Z
dateModified: 2026-10-10T00:00:00.000Z
author: palaniappan-p
category: Cost Optimization & FinOps
tags: amazon-bedrock, knowledge-bases, gmail, google-drive, s3-vectors, cost-optimization, rag
---

# Bedrock Knowledge Base Cost for Gmail and Google Drive

> Checked 10 October 2026: managed Bedrock knowledge base storage is $5 per GB of raw data a month. One thousand GB is $5,000 before retrieval. An e-commerce workspace of terabytes of Drive files and millions of Gmail messages reached hundreds of dollars a day after ingestion started.

On 10 October 2026 we re-checked the [Amazon Bedrock pricing page](https://aws.amazon.com/bedrock/pricing/) before writing another ingestion job. Managed Knowledge Base storage is **$5.00 per GB of raw data per month**. One thousand of those gigabytes is **$5,000 a month**. Dividing by 30.4 days, that is about **$164 a day**, before a single Retrieve call.

An e-commerce operation keeps terabytes of Google Drive files and millions of Gmail messages. We connected that workspace to a Bedrock knowledge base and started ingestion. AWS spend moved into the hundreds of dollars a day. The service-by-service split is still open. The meter above is large enough to produce a bill of that shape once the raw corpus is multi-terabyte. Live search of Gmail and Drive then missed some specialized questions that needed several historical documents at once.

> **From a real engagement** — E-commerce operations knowledge. Terabytes of Google Drive. Millions of Gmail messages. AWS spend climbed to hundreds of dollars a day once knowledge base ingestion started. The line items are still open. We stopped treating the whole workspace as the index.

> **What broke** — During ingestion. The daily bill entered the hundreds before anyone had a diagnosis. Live Gmail and Drive search then failed on specialized questions that span several older documents. Detection was the bill, then those missed answers. Recovery was to halt full-corpus indexing and split the path: live search for the archive, a curated index for the question classes live search cannot carry. The split below is the design we recommend. We have not production-validated it.

> **Reproduce this** — The worksheet prints the storage illustration from the rates in this post. It does not call AWS and it does not score retrieval. `python3 examples/architecture-blog-2026/gmail-drive-selective-kb/cost_worksheet.py` prints `1,000 GB x $5.00 = $5,000.00 per month` and `1,024 GB x $5.00 = $5,120.00 per month`. Folder: [gmail-drive-selective-kb](/examples/architecture-blog-2026/gmail-drive-selective-kb/README.md).

We recommend live Gmail and Drive search as the default, and a customer-managed Bedrock Knowledge Base on Amazon S3 Vectors for the question classes that fail that search on a shared scorecard. Use the managed Google Drive connector only when you have measured the raw gigabytes and those gigabytes times $5 fit the budget you will actually watch.

## A terabyte on the managed meter

[Managed Knowledge Bases](https://aws.amazon.com/bedrock/pricing/) and customer-managed Knowledge Bases are different products. The pricing page, checked 10 October 2026, bills the managed product like this:

| Meter | Rate on the pricing page |
| --- | --- |
| Index storage | $5.00 per GB of raw data per month |
| Standard Retrieve | $1.00 per 1,000 API calls |
| Agentic retrieval | $4.00 per 1,000 Agentic Retrieve calls, plus $1.00 per 1,000 underlying Retrieve calls |
| Managed parser, managed embedding model, managed reranker | Included |
| Your own embedding model, reranker, or planning model | That model's Bedrock price, on top |

The page does not say whether the storage GB is 10^9 bytes or 2^30 bytes. [Amazon S3's pricing page](https://aws.amazon.com/s3/pricing/) does say it, for S3: a GB is 2^30 bytes, and 1 TB is 1,024 GB. Those sentences belong to S3. Apply them to the Bedrock meter only when the billed quantity is the binary one.

The illustration, and only an illustration:

- If the managed meter reports 1,000 GB of raw data, storage is 1,000 × $5 = **$5,000 per month**.
- If it reports 1,024 GB, storage is 1,024 × $5 = **$5,120 per month**.

Each extra 1,000 GB on that meter is another $5,000 a month, about $164 a day at a 30.4-day divisor. A raw index of a few thousand gigabytes lands in the hundreds of dollars a day on storage alone. That is a candidate explanation for the bill we saw. It is not a diagnosis. We have not read the line items.

Customer-managed Knowledge Bases have no separate hourly knowledge base fee on that same pricing page. You pay embeddings, the vector store, optional parsing, optional rerank, and the model that writes the answer. The [9 October cost post](/blog/bedrock-knowledge-base-cost-optimization-2026/) prices one 5,000-document corpus on that path. OpenSearch Serverless Classic in that post sits on a floor of $350.40 a month (2 OCUs × 730 hours × $0.24). A floor of that size does not, by itself, explain hundreds of dollars a day. A scaled OCU fleet, a full re-embed, Bedrock Data Automation at $0.010 per page, or model tokens can. On the managed product the parser and the managed embedding model are included, so a parsing spike is a customer-managed story, not the managed-storage story.

A [custom managed connector](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-ds-custom.html) can take Gmail we normalize ourselves. It still sits on the $5 per GB raw meter. Custom ingest is how you get mail into a managed knowledge base. It is not how you escape the storage price.

## What the first ingestion showed

Three facts, kept separate.

**Documented rates.** Managed storage at $5 per GB-month. Retrieve at $1 per 1,000 calls. Agentic retrieval at $4 per 1,000 calls plus the underlying Retrieve meter. S3 Vectors, in the pricing-page example checked the same day, at $0.06 per GB-month of logical vector storage, $0.20 per GB uploaded, and $2.50 per million query API requests. Data processed in that example is $0.004 per TB for the first 100,000 vectors in the index and $0.002 per TB after that. Returned data in the example has a 500 KB free allowance per query. On 16 June 2026, AWS cut data-processed charges by up to 80% for indexes over 10 million vectors. Confirm the regional row before you lock a budget. The [site calculator](/tools/aws-bedrock-knowledge-base-pricing-calculator/) still models S3 Vectors at an illustrative $1 per GB-month and $0.05 per 1,000 queries, dated in the 9 October post. Those two figures are not the pricing-page rates. Use the calculator for the self-managed 5,000-document comparison it was built for. Use the pricing page for S3 Vectors.

**What we observed.** After the data sources were connected and ingestion started, daily AWS cost entered the hundreds. Live search of Gmail and Drive answered many specific lookups and missed some specialized questions. We do not have a percentage, a latency, or a saved-dollar figure for either observation.

**What we have not shown.** Which line item produced the spike. Whether the index was managed raw storage, OpenSearch capacity, embedding tokens, parsing pages, or model calls. Whether the hybrid design below would have changed the answers. Treat every dollar in the next sections as list-price arithmetic, not as a result from this engagement.

## Live search carries the archive

Gmail already has a search language. [`users.messages.list`](https://developers.google.com/workspace/gmail/api/reference/rest/v1/users.messages/list) accepts `q` in the same form as the Gmail search box, returns at most 500 messages per page, and requires a scope that can see the body if you want `q`. The `gmail.metadata` scope cannot use `q`. [`gmail.readonly`](https://developers.google.com/workspace/gmail/api/auth/scopes) is the scope that matches a read-only agent. Then fetch the thread, not every quoted copy of the same conversation.

Drive's [`files.list`](https://developers.google.com/workspace/drive/api/reference/rest/v3/files/list) accepts `q` for MIME type, modified time, parents, and full text. Page size tops out at 1,000. If `incompleteSearch` is true, the call did not cover every corpus. Narrow to `user` or a single shared drive and search again. `drive.readonly` can see the archive the user can see. `drive.file` sees files the app created or the user opened with the app, which will not see an existing shared drive.

Cap the pages. A question that needs the 40th page of hits is a sign the query is too wide, not a sign you should embed the mailbox. Keep the message id, thread id, and Drive `webViewLink` on the answer so a person can open the source.

Cache a repeated query on the user, the query string, and the scopes, with a short TTL. Drop the entry when [`users.history.list`](https://developers.google.com/workspace/gmail/api/reference/rest/v1/users.history/list) or the Drive change log moves. Do not reuse one user's cache for another user.

Live search is the right default when the question names a sender, a label, a filename, a folder, or a date. It is the wrong tool when the question needs a meaning that is spread across several documents and no keyword you can write down. That is what our experiments hit. Keyword search was weak on those questions. That is an observation about those questions, not a claim that Google's search is generally inaccurate.

## Claude's connectors stay in the Claude app

[Claude's Google Workspace connectors](https://support.claude.com/en/articles/10166901-use-google-workspace-connectors) let Claude search Gmail, Calendar, and Drive from a chat, with citations back to the originals. Gmail attachment content is metadata only. Team and Enterprise owners have to enable the connectors before a member can authenticate. The useful lesson is the retrieval shape: ask, then fetch the minimum, under the person's existing Google permissions.

That configuration lives in Claude. A production agent on Amazon Bedrock does not inherit it.

The [Anthropic API MCP connector](https://platform.claude.com/docs/en/agents-and-tools/mcp-connector) is a third thing. The Messages API, with beta header `mcp-client-2025-11-20`, calls tools on a public HTTP MCP server. Local stdio servers are out. It does not reuse the OAuth grant a person clicked inside the Claude app.

The Bedrock agent needs its own grant. Per-user OAuth with `gmail.readonly` and `drive.readonly` matches an agent that answers as that person. Domain-wide delegation is an admin decision, with an audit trail, for a workspace-wide job. One shared admin token that can read every mailbox is the wrong default. Store the credential in Secrets Manager. Log the tool name, the user, and the source ids. Do not log message bodies.

## Route the question before you retrieve

An LLM call that only picks a retrieval path is still an LLM call. Start with rules.

Python 3.12. No AWS SDK. Replace the topic set with the classes your scorecard actually failed.

```python
def route(question: str, curated_topics: set[str]) -> str:
    text = question.casefold()
    if any(token in text for token in ("latest", "from:", "filename:")):
        return "live"
    if any(topic in text for topic in curated_topics):
        return "curated"
    if "policy" in text and "today" in text:
        return "both"
    return "live"
```

`live` calls Gmail or Drive. `curated` calls Knowledge Bases Retrieve on the small index. `both` calls live search and the index, then the model has to say which source is newer. Send the question to a classifier model only when the rules return an unknown class you have already listed. A wrong route to the index is cheaper than indexing the archive, and more expensive than one Gmail query. Log the route so you can see which rule fired.

```mermaid
flowchart LR
  question[Question] --> router[Deterministic router]
  router --> liveMail[Gmail API]
  router --> liveDrive[Drive API]
  router --> kb[Curated knowledge base]
  liveMail --> answer[Cited answer or refusal]
  liveDrive --> answer
  kb --> answer
  select[Approved folders labels threads] --> s3[S3 curated corpus]
  s3 --> kb
  kb --> vectors[S3 Vectors]
```

The draw.io file has the same path with service icons: [Gmail and Drive hybrid retrieval](/examples/architecture-blog-2026/gmail-drive-selective-kb/gmail-drive-hybrid-knowledge-base.drawio).

Live search and the curated index cover different content. A second retriever on the same corpus is a defect: the agent then reconciles two rankings, and a permission filter on only one of them is a hole. One corpus, one retrieval API, and the permission filter on that API.

## What belongs in the curated corpus

Index a document when the scorecard shows live search missed a repeated question and the document is one of the sources that would have answered it. Candidates, each of which still has to earn its place:

- Canonical files in named Drive folders: policies, vendor terms, runbooks.
- Threads with a label, a sender, and a date window that the scorecard keeps failing.
- A normalized summary only when the summary stores the source link, the timestamp, and the fact that it is a summary.

Skip the rest of the mailbox. Quoted reply chains, signatures, and duplicate copies of the same thread inflate raw gigabytes and repeat stale answers. Keep one thread record, the new message, and a hash of the normalized text so a re-sync does not re-embed an unchanged body. If two sources disagree, keep both and the dates. The model should cite the conflict, not average it.

Gmail has no managed connector. The path for mail is: our job selects the thread, writes a normalized object to S3 with metadata (`tenant`, `source`, `messageId`, `threadId`, `modifiedTime`), and a customer-managed knowledge base syncs that prefix into S3 Vectors. Drive can take the same path, which is what we recommend when Gmail and Drive must share one index and one permission filter.

The managed Drive connector is the other Drive option. Checked on the [connector page](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-ds-googledrive-connect.html) on 10 October 2026, you can limit shared drives, MIME types, a modified-date window, and file size (default 500 MB). Folder and file ids are available with OAuth, not with a service account. Document ACLs require a service account, and `aclEnabled` cannot be changed after the data source is created. If you omit it, a service account defaults to on and OAuth defaults to off. A redacted scope is in [drive-filter.example.json](/examples/architecture-blog-2026/gmail-drive-selective-kb/drive-filter.example.json). The [connect page](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-managed-connect-ds.html) lists connector types `S3`, `ONEDRIVE`, `CONFLUENCE`, `SHAREPOINT`, `WEB_CRAWLER`, and `GOOGLE_DRIVE`. The create page also names Box, Salesforce, ServiceNow, and Zendesk. Gmail is on neither page.

Pay the managed connector when three things are true at once: the bytes you will actually crawl times $5 per month are acceptable, the files live in Drive, and you want included hybrid search plus document ACLs. Hybrid search matters because [S3 Vectors with Knowledge Bases is semantic search only](https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-vectors-bedrock-kb.html). A filename lookup that semantic search muffs can stay on the managed Drive index, or on live `files.list`, which is usually the cheaper of those two.

S3 Vectors storage is not the same unit as managed raw storage. The $0.06 per GB-month example price is logical vector storage: dimensions, metadata, and key. A terabyte of PDFs does not become a terabyte of vectors, and this post does not pretend to convert one into the other. Measure chunk count and vector bytes on the curated set, then price that. The 17 July 2025 [AWS walkthrough](https://aws.amazon.com/blogs/machine-learning/building-cost-effective-rag-applications-with-amazon-bedrock-knowledge-bases-and-amazon-s3-vectors/) is how you attach S3 Vectors to a knowledge base. Its metadata sizes are from the preview. The current integration page caps Knowledge Bases custom metadata at 1 KB and 35 keys per vector. Put `tenant`, `source`, and `modifiedTime` in those keys. A long ACL will not fit. Enforce the ACL before the chunk is written, or keep that corpus on the managed Drive connector where document ACLs are a product feature.

Query latency on that integration page is sub-second when cold and as low as 100 milliseconds when warm. If the agent loop needs the tighter budget already described in the [S3 Vectors post](/blog/amazon-s3-vectors-native-vector-storage/), this store is the wrong row. That post compares S3 Vectors with OpenSearch Serverless and MemoryDB.

## Retrieval quality is its own budget

RAG is not the prize for losing at keyword search. It is a second index with its own failure modes. Use a technique when the scorecard says the previous pass missed, and skip it when the first pass already cited the right source.

Query rewriting costs a model call. Skip it when the person already typed a Gmail or Drive query. Metadata filters are cheap relative to a second model call, and on S3 Vectors they have to fit the 1 KB and 35-key cap. Ask for a small candidate set first. The 9 October post prices Amazon Rerank 1.0 at $1 per 1,000 queries in the site calculator, and Cohere Rerank 3.5 at $2 per 1,000 queries on the Bedrock pricing page, with at most 100 chunks in a Cohere query. Turn rerank on for the question class that missed, not for every call. The managed knowledge base includes its own reranker. A customer-managed index does not.

Group Gmail hits by thread before they reach the model. Prefer the newer source when two chunks disagree, and show both dates when the disagreement is the answer. If the curated index returns nothing useful, run the live query before you invent a span. If live search also returns nothing the question asked for, say so and ask a narrower question. A citation the reader can open is part of the answer. A paragraph with no message id and no file link is not.

## Four strategies and one storage illustration

| Strategy | Build | Ongoing cost | Retrieval | Freshness | Operations | Security |
| --- | --- | --- | --- | --- | --- | --- |
| Whole workspace in a managed knowledge base | Connector setup, then a long first sync | $5 per raw GB-month plus Retrieve. 1,000 GB is $5,000 a month in the illustration | Hybrid search included. Quality still depends on what you crawled | Sync on demand, daily, weekly, or monthly | Least code. Hardest bill to cap once the crawl is wide | Drive ACLs if you set them at create time, with a service account. No Gmail connector |
| Live Gmail and Drive only | OAuth, two tools, retries, a page cap | Google API quota and a short cache. No vector store | Strong on lookups you can name. Weak on the specialized questions we already missed | As fresh as the API | You own rate limits, 404 history ids, and `incompleteSearch` | The user's own permissions, if the token is that user's |
| Scoped managed Drive connector, plus live retrieval | OAuth or service account, folder and MIME filters, live tools for everything else | $5 per GB of the scoped Drive set, plus live-call cost. Gmail stays off the meter | Hybrid search on the Drive slice. Mail stays keyword search | Drive follows the sync schedule. Mail is live | Watch the raw GB of the scope. Folder ids are OAuth-only | Drive ACLs available. Mail follows the live token |
| Curated S3 corpus, customer-managed Knowledge Base, S3 Vectors, plus live retrieval | Selection job, S3 layout, knowledge base, scorecard | Embeddings, S3 object storage, S3 Vectors on vector bytes, optional rerank, model tokens. Do not reuse the $5,000 figure for this row | Semantic search on the curated set. Live tools for the rest. No hybrid search inside S3 Vectors | Live path is current. Index moves when the job writes and the sync finishes | The most code. The clearest place to refuse a document | Filter before write. Tenant id on every Retrieve. Separate prefix per tenant |

We recommend the fourth row for this workload: terabytes of files, millions of messages, Gmail in scope, and a handful of question classes that keyword search missed. The third row wins when the curated set is Drive-only, small enough that $5 per GB is dull, and you want ACLs and hybrid search without running the embedding pipeline. The second row wins if the scorecard never shows a class that live search misses. The first row is how the daily bill got large.

## Permissions, sync, and a stop switch

Gmail push notifications come from [`users.watch`](https://developers.google.com/workspace/gmail/api/guides/push). Google's guide tells you to call `watch` again at least every 7 days, and to call `history.list` when a notification does not show up. A `historyId` that returns HTTP 404 is stale. Full sync that mailbox, then store the new id. Drive's [change log](https://developers.google.com/workspace/drive/api/guides/change-overview) is not a copy of every edit: an ACL change shows up for the owner, for service accounts on the ACL, and for the users the change affects. Take a start page token, poll `changes.list` on a schedule, and reconcile. Push is a hint. The schedule is the backstop.

[Knowledge base sync is incremental](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-data-source-sync-ingest.html). Unchanged documents are skipped. Changed documents are parsed, chunked, embedded, and indexed again. Deleted documents leave the vector store. You can stop a running job with `StopIngestionJob`. Managed connectors can also sync daily, weekly, or monthly. Deletion protection, default threshold 15 percent, skips the delete phase when a sync would remove more than that share of the index. That saves you from a bad crawl. It also leaves stale vectors in place after a real bulk delete until you raise the threshold on purpose. The custom connector does not support deletion protection. You delete those documents yourself.

AWS CLI 2, `bedrock-agent`. The job must already be running. This call is not attached to a budget.

```bash
aws bedrock-agent stop-ingestion-job \
  --knowledge-base-id "$KB_ID" \
  --data-source-id "$DS_ID" \
  --ingestion-job-id "$JOB_ID"
```

Make the writer idempotent. Content hash in the object key. A dead-letter queue for Google API errors you have retried. A per-job cap on documents and bytes, checked before `StartIngestionJob`. CloudWatch on pages fetched, documents written, documents refused, ingestion job state, and Retrieve errors. AWS Budgets and Cost Anomaly Detection on the Bedrock and S3 lines, with a named human. A budget notification does not halt the API. The cap in the job, and `StopIngestionJob`, are the halt.

Encrypt the curated bucket with a customer-managed key. Keep Google credentials in Secrets Manager. One S3 prefix and one knowledge base per tenant, or a `tenant` metadata filter on every Retrieve and nowhere else. S3 Vectors will not hold a full permission list inside 1 KB of filterable metadata. If the permission model is "whoever can open the file in Drive," prefer live search or the managed Drive ACL over a copied ACL in the vector index.

## Six stages before the corpus grows

The sizes and the pass marks come from your bill and your questions. This post does not have a universal document count.

**Stage 1. Establish the baseline.** Read the AWS bill. Separate Bedrock, S3, S3 Vectors, and OpenSearch. Note the indexed byte count if the console shows one. Write down the questions staff actually ask, ranked by how often they block a decision.

**Stage 2. Build an evaluation set.** For each question, record the class (live, curated, or both), the evidence that would count, and who is allowed to see it. Use the [blank scorecard](/examples/architecture-blog-2026/gmail-drive-selective-kb/evaluation-scorecard.md). Keep customer text out of the ticket.

**Stage 3. Compare the paths.** Run live search, the index you have now, and a tiny curated index against the same questions. Record correctness, whether the citation opens the right source, latency, tool failures, permission misses, and cost per answered question.

**Stage 4. Pilot the curated set.** Index one folder or one label that Stage 3 showed was worth it. Add the next category only when the sheet moves.

**Stage 5. Add the controls.** Incremental Gmail history and Drive changes, a reconcile job, deletion on the index, retries, the dead-letter queue, the byte cap, CloudWatch, and the budget alarm wired to a person who can stop the job.

**Stage 6. Expand when the sheet says so.** A new category joins the index when accuracy on that class improves enough to justify the raw gigabytes or the vector bytes you measured. Otherwise it stays on live search.

## What to Do This Week

1. Open Cost Explorer for the days ingestion ran. Group Bedrock by feature if the window starts on or after 1 September 2026, then read S3, S3 Vectors, and OpenSearch for the same days. The [9 October post](/blog/bedrock-knowledge-base-cost-optimization-2026/) has the `PRODUCT_ATTRIBUTE` calls.
2. If a knowledge base sync is still running and you cannot explain the raw gigabytes, stop it with `StopIngestionJob`.
3. Copy the [scorecard](/examples/architecture-blog-2026/gmail-drive-selective-kb/evaluation-scorecard.md) and fill ten real questions. No customer text in the copy you commit.
4. Run `python3 examples/architecture-blog-2026/gmail-drive-selective-kb/cost_worksheet.py` and write your own indexed-GB figure next to the $5,000 illustration.
5. Put a byte cap in front of the next `StartIngestionJob`. Attach a budget alert to a person, and do not describe the alert as a cap.

## What This Post Doesn't Cover

Calendar. OCR of Gmail attachments. Amazon Bedrock AgentCore Memory, which is session state, not this corpus. A measured retrieval score for the hybrid design, because we have not run it in production. The line items of the bill that started this work. A universal corpus size. Kendra's Gmail connector as a path forward: Kendra is in maintenance for new customers after 30 July 2026, which the [lifecycle roundup](/blog/aws-service-lifecycle-updates-june-2026/) already records.

The [RAG pipeline guide](/blog/how-to-build-rag-pipeline-amazon-bedrock-knowledge-bases/) is the setup walkthrough. This page is the decision about what you are willing to store.

If you want a second set of eyes on the bill and the scorecard before the next sync, [talk to us about the retrieval path](/contact-us/?focus=ai-agents). The business entry for a knowledge agent is [the knowledge agent page](/ai-agents/knowledge/). The AWS build is [Amazon Bedrock](/services/aws-bedrock/). Price a self-managed corpus on the [knowledge base calculator](/tools/aws-bedrock-knowledge-base-pricing-calculator/), then check the S3 Vectors row against the pricing page rather than against the calculator's illustrative vector rates.

## FAQ

### When should we not index Gmail into a knowledge base?
Leave the mailbox on live Gmail API search when the question names a sender, label, date, or thread, and when you have not shown that semantic search across several historical threads beats that query. Millions of messages on a managed knowledge base are billed as raw gigabytes at $5 per GB-month. Index a thread only after the evaluation sheet shows that live search missed it.

### What goes wrong if a budget alert is the only spending stop?
AWS Budgets and Cost Anomaly Detection notify people. They do not cancel a running ingestion job. StopIngestionJob stops a job that is already in progress, and only if something calls it. Put a document count and a byte cap in the ingestion function before StartIngestionJob, and page a human who can call StopIngestionJob when the alert fires.

### Can an Amazon Bedrock agent call Claude's Google Workspace connectors?
No. Those connectors run inside Claude and Claude Desktop, on the Google account a person connected there. The Anthropic API MCP connector is a different feature: the Messages API calls a public HTTP MCP server. A Bedrock agent needs its own Google OAuth or domain-wide delegation and its own Gmail and Drive tools.

### Does Bedrock have a native Gmail knowledge base connector?
Not on the connector pages checked 10 October 2026. Managed knowledge bases list Google Drive, plus S3, SharePoint, OneDrive, Confluence, a web crawler, and a custom connector. The create page also names Box, Salesforce, ServiceNow, and Zendesk. Gmail is on neither list. Gmail has to be a custom integration. Amazon Kendra has a Gmail connector, and Kendra is in maintenance for new customers after 30 July 2026.

### When is S3 Vectors the wrong store for this corpus?
When the tool loop needs the latency already sent to MemoryDB or OpenSearch on the S3 Vectors post, or when you need hybrid keyword and vector search in one engine. S3 Vectors with Knowledge Bases is semantic search only. A small Drive-only set that needs document ACLs and included hybrid search can stay on a scoped managed Google Drive connector instead.

### When is the managed Google Drive connector the better bill?
When the content is Drive-native, you have measured the raw gigabytes, and those gigabytes times $5 per month fit the budget you will watch. The managed path includes the parser, the managed embedding model, the managed reranker, and hybrid search. It is a poor default for a multi-terabyte workspace, and it still does not ingest Gmail.

---

*Source: https://www.factualminds.com/blog/bedrock-knowledge-base-cost-gmail-google-drive/*
