---
title: CDC into S3 Tables: Zero-ETL or DMS plus MERGE
description: A lab query still returned a customer the source deleted. Zero-ETL can land DynamoDB and listed SaaS sources in S3 Tables. A self-managed database cannot. Firehose database CDC was removed on 24 September 2025.
url: https://www.factualminds.com/blog/cdc-into-s3-tables-lakehouse/
datePublished: 2026-10-04T00:00:00.000Z
dateModified: 2026-10-04T00:00:00.000Z
author: palaniappan-p
category: Data & Analytics
tags: aws, amazon-s3, aws-glue, aws-dms, apache-iceberg, cdc
---

# CDC into S3 Tables: Zero-ETL or DMS plus MERGE

> A lab query still returned a customer the source deleted. Zero-ETL can land DynamoDB and listed SaaS sources in S3 Tables. A self-managed database cannot. Firehose database CDC was removed on 24 September 2025.

A query against the lab S3 Table still returned a customer the source had deleted that morning. DMS had written operation `D`. The sink appended the file as another insert. Compaction was the proposed fix. Compaction rewrites files. It does not know the business key.

Checked on **4 October 2026**. [Glue Zero-ETL](https://docs.aws.amazon.com/glue/latest/dg/zero-etl-using.html) can target **Amazon S3 Tables** through the SageMaker lakehouse from DynamoDB, Oracle at AWS, and the SaaS sources on that page. The same page says self-managed Oracle, SQL Server, MySQL, and PostgreSQL replicate **only to Redshift**. [DMS can write CDC to S3](https://docs.aws.amazon.com/dms/latest/userguide/CHAP_Target.S3.html), including a transaction file whose first field is the operation (`I`, `U`, or `D`). [Firehose's database source was removed on 24 September 2025](https://docs.aws.amazon.com/firehose/latest/dev/history.html). The mover choice itself is the [ingestion decision](/blog/glue-zero-etl-vs-dms-vs-firehose-vs-appflow/). This page is only the two legal shapes into the lake, and what deletes do.

![Lab figure of two legal CDC shapes into S3 Tables, with a self-managed database blocked from Zero-ETL. Not a customer account.](../../assets/images/blog/cdc-into-s3-tables-lakehouse-01.webp)

*Lab figure, 4 October 2026. Not a customer account. Shape A only when Glue allows the pair. Otherwise DMS to S3, then a MERGE that applies the delete.*

Teams finish a database move and then build a second, custom capture path into the lake. Or they full-refresh a large table every night because the lake project never had an incremental design. Both are the same miss. [S3 Tables](/blog/aws-modern-data-lake-s3-tables-iceberg-reference-architecture-2026/) and the [lakehouse pattern](/patterns/lakehouse-on-aws/) already describe the table. The [DMS to Aurora playbook](/blog/aws-database-migration-dms-aurora-playbook-2026/) already describes a cutover into Aurora. Neither is this path.

> **Reproduce this** — Copy the [two-shape checklist](/examples/architecture-blog-2026/cdc-s3-tables/two-shape-checklist.md). One source table, one shape, one answer for deletes, one answer for an added column.

## Shape A — Zero-ETL straight into S3 Tables

Use this when the source is on the Glue list **and** the target restriction allows S3 Tables.

That is DynamoDB, Oracle Database@AWS, or a SaaS application the Zero-ETL page names (Salesforce, SAP OData, ServiceNow, Zendesk, Zoho CRM, and the ads and marketing sources on that page — re-read it). The integration lands the data through the SageMaker lakehouse into S3 Tables. You do not operate a replication instance for the copy, and you do not write a MERGE job to apply ordinary updates from that source.

Do not use Shape A for a self-managed database. The docs forbid the target. Do not use it because a workshop used PostgreSQL. If the console will not create the integration, believe the console and switch to Shape B.

The trade-off: you get the pair AWS operates, and you do not get an arbitrary transform inside the mover. Transform after the table exists, in a job that reads the S3 Table. [Glue 5 and Iceberg](/blog/aws-glue-5-apache-iceberg-modern-etl/) is where that job's table format belongs. This article does not repeat it.

## Shape B — DMS CDC to S3, then MERGE

Use this when the source is a database and Shape A is illegal or is not the target you need. RDS and Aurora included: they are not the self-managed Zero-ETL class, and they are also not, by themselves, a reason to invent a third tool. If a future Glue page adds your engine as an S3 Tables source, revisit Shape A. Until that sentence exists, DMS is the capture.

DMS full load plus ongoing replication writes changes to S3. By default, each table's changes land in that table's prefix, and those files are **not** ordered by transaction. If you need transaction order, the S3 target settings can write transaction files that list every row as `operation,table_name,database_schema_name,...` with `I`, `U`, or `D`. Pick one of those layouts and make the sink match it. A sink written for per-table files will misread a transaction file, and the reverse.

The commit into S3 Tables is a Glue job that **MERGEs** into the Iceberg table. Inserts and updates are the easy half. A delete is the point of CDC. The job has to apply `D` as a delete (an Iceberg MERGE matched delete, or an equality delete). An append of every file DMS wrote will keep the deleted row forever and call it history.

We recommend Shape A over Shape B whenever the pair is legal, because you are not operating the MERGE. We recommend Shape B over a nightly full refresh when the table is large enough that rewriting it is the cost. The trade-off of Shape B is a job you must make idempotent: a rerun that inserts the same change twice is a duplicate, not a retry. Glue bookmarks and Iceberg commit conflicts are separate articles. Do not pretend this page specified them.

## Deletes and schema changes

These are the two tests. A pipeline that only loads inserts is not CDC.

**Deletes.** After you delete a row in the source, the same key must disappear from the query against the S3 Table, or you have documented why the lake keeps it. "We will compact later" does not delete a row. Compaction rewrites files. It does not know which business key the source removed. The [modern data lake post](/blog/aws-modern-data-lake-s3-tables-iceberg-reference-architecture-2026/) owns compaction. Leave it there.

**Schema changes.** Add a column in the source. Confirm the next commit either adds that column to the Iceberg schema or fails on purpose. A silent drop means the lake and the source have diverged and Athena will not tell you in the task status. One column, in a non-production table, is the test. A full reload is the wrong test: it hides a sink that cannot evolve.

Do not run Shape A and Shape B on the same table. Two writers, one Iceberg table, and the row count drifts. The ingestion article already says to stop the second mover.

## Where this fails

![Lab figure showing an append-only sink that kept a deleted key, beside a MERGE that removed it from the S3 Table. Not a customer account.](../../assets/images/blog/cdc-into-s3-tables-lakehouse-02.webp)

*Lab figure, 4 October 2026. Not a customer account. The check is the business key: gone in the source, gone in the query. Compaction is not that check.*

DMS is healthy. The S3 prefix grows. A query still returns a customer the source deleted yesterday. Detection is a key that exists in the lake and not in the source, not a replication lag of zero. The sink appended `D` as another row, or it ignored the operation field. The fix is a MERGE that deletes, then a rerun that does not insert the deleted key again. A second nightly copy of the whole table will make the query look right and will hide the sink.

The other failure is Shape A specified for self-managed PostgreSQL "into S3 Tables." The Glue page's note is the spec: Redshift only. The integration should not be created. If a design review approved it anyway, the review skipped the source class. Send that source to Shape B, or to Redshift via the [Redshift playbook](/blog/aws-redshift-data-warehouse-modernization-playbook-2026/) if the warehouse was the real target.

## What to Do This Week

1. Write down the source class: DynamoDB, Oracle at AWS, listed SaaS, or a database outside that list.
2. If the class is allowed into S3 Tables, choose Shape A and do not also create a DMS task.
3. Otherwise choose Shape B. Decide per-table files or transaction files before anyone writes the job.
4. Test one delete and one added column against a non-production table. Keep the results with the checklist.
5. Leave compaction, Aurora cutover, and the choice among Firehose and AppFlow on their own URLs. If the path is the platform, [data analytics](/services/aws-data-analytics/) is the commercial next step.

## What This Post Doesn't Cover

LOB lag, Aurora cutover weekends, and DMS task settings stay on the playbook. Iceberg compaction pricing and catalog choice stay on the S3 Tables reference. Glue 5 DDL stays on that post. High-volume MERGE tuning, Glue bookmarks, and commit conflicts are not written here. Firehose into Iceberg as a stream destination is not database CDC. This page does not estimate a terabyte or a nightly refresh cost. Re-check the Glue Zero-ETL source list before you lock Shape A.

## FAQ

### How should operational changes land in S3 Tables?
Two shapes. Shape A: a Glue Zero-ETL pair the docs allow straight into S3 Tables (DynamoDB, Oracle at AWS, or a listed SaaS source, through the SageMaker lakehouse). Shape B: DMS CDC to S3, then a Glue MERGE into the S3 Table, for a database Zero-ETL will not send there. Do not add a third pipeline.

### When should we not use Zero-ETL into S3 Tables?
When the source is self-managed Oracle, SQL Server, MySQL, or PostgreSQL. Glue Zero-ETL replicates those sources only to Amazon Redshift. Forcing that pair at S3 Tables is not a configuration tweak. Use Shape B, or land in Redshift and stop.

### What could go wrong with deletes?
DMS can write the operation on each changed row, including a delete. An append-only sink ignores that operation. The source row is gone and the lake still returns it. The MERGE, or an Iceberg equality delete, has to apply the delete. A nightly full copy hides the bug and rewrites the table.

### What could go wrong with a schema change?
A new column in the CDC file does not alter the S3 Table by itself. The next MERGE either drops the field or fails, depending on how the job is written. Evolve the Iceberg schema in the job that commits, and test one added column before you call the sink done. This article does not teach compaction.

### Is Firehose the CDC path into Iceberg?
No. Database as a Firehose source was removed on 24 September 2025. Firehose can deliver streams to Iceberg tables. That is a different ingest. Database changes use Shape A or Shape B.

### Where is the Aurora cutover, and where is compaction?
Aurora cutover, including LOB lag and task checklists, stays on the DMS playbook. Compaction and the S3 Tables reference architecture stay on the modern data lake post. MERGE syntax and Glue 5 table patterns stay on the Glue 5 Iceberg post. This page only chooses the CDC path and the delete rule.

---

*Source: https://www.factualminds.com/blog/cdc-into-s3-tables-lakehouse/*
