goutamadwant opened a new pull request, #12587:
URL: https://github.com/apache/seatunnel/pull/12587
### Purpose of this pull request
Fix #12585.
PostgreSQL 13+ can invalidate a replication slot (`max_slot_wal_keep_size`,
`idle_replication_slot_timeout` on 18). Postgres-CDC passed such a slot to
Debezium, which then either looped on the slot state for about 30 minutes
(`wal_removed`) or failed after a minute of stream retries (`idle_timeout`), on
every pipeline restore and with no hint about the slot.
`PostgresSourceFetchTaskContext.configure()` now checks
`pg_replication_slots` before Debezium uses the slot, when the job streams:
- The slot is invalidated if `invalidation_reason` (PostgreSQL 17+) is set,
`conflicting` (16+) is true, or `wal_status` (13+) is `lost`. Only columns the
server returns are read, so older servers are unaffected.
- The reader fails with the new `POSTGRES-04`, which names the reason and
how to recover (drop the slot, restart without restoring state).
- Snapshot-only jobs skip the check. If the query itself fails, the reader
logs a warning and continues as before.
### Does this PR introduce _any_ user-facing change?
Yes. A job on an invalidated slot now fails quickly with `POSTGRES-04`
instead of hanging or failing with a Debezium error. Documented in the
PostgreSQL-CDC FAQ (en, zh).
| Case (default `job.retry.times = 3`) | before | after |
|---|---|---|
| PG 18 `idle_timeout`, restore | FAILED after ~4 min 20 s | FAILED in 13 s,
`POSTGRES-04` |
| PG 17 `wal_removed`, restore | RUNNING, slot loop 900 × 2 s per attempt |
FAILED in 13 s, `POSTGRES-04` |
| PG 17 slot invalidated while running | hung after the pipeline restore |
FAILED in 15 s |
| Healthy slot restore, fresh start, snapshot-only | works | unchanged |
### How was this patch tested?
- Added tests to `PostgresUtilsTest` (PostgreSQL 17+ reasons, `conflicting`,
`wal_status = lost`, servers without these columns) and a new
`PostgresSourceFetchTaskContextTest` (snapshot-only skips the query,
invalidated slot throws, healthy slot passes, query error rolls back). `./mvnw
-pl seatunnel-connectors-v2/connector-cdc/connector-cdc-postgres test` passes
on JDK 8 and JDK 11.
- `PostgresCDCIT` passes locally.
- Manual runs on PostgreSQL 18.6 and 17.9 for the cases in the table, plus
committed-offset mode and parallelism 2 on an invalidated slot.
### Check list
* [ ] If any new Jar binary package adding in your PR, please add License
Notice according
[New License
Guide](https://github.com/apache/seatunnel/blob/dev/docs/en/developer/new-license.md)
* [x] If necessary, please update the documentation to describe the new
feature. https://github.com/apache/seatunnel/tree/dev/docs
* [ ] If necessary, please update `incompatible-changes.md` to describe the
incompatibility caused by this PR.
* [ ] If you are contributing the connector code, please check that the
following files are updated:
1. Update
[plugin-mapping.properties](https://github.com/apache/seatunnel/blob/dev/plugin-mapping.properties)
and add new connector information in it
2. Update the pom file of
[seatunnel-dist](https://github.com/apache/seatunnel/blob/dev/seatunnel-dist/pom.xml)
3. Add ci label in
[label-scope-conf](https://github.com/apache/seatunnel/blob/dev/.github/workflows/labeler/label-scope-conf.yml)
4. Add e2e testcase in
[seatunnel-e2e](https://github.com/apache/seatunnel/tree/dev/seatunnel-e2e/seatunnel-connector-v2-e2e/)
5. Update connector
[plugin_config](https://github.com/apache/seatunnel/blob/dev/config/plugin_config)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]