DanielLeens commented on issue #12267: URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5749724877
Thanks for supplying the complete catalog result. This confirms the diagnosis. There is one physical `testdb.test_a` relation in `pg_class`/`pg_namespace` (OID `334287`), while HighGo's `information_schema.tables` returns six identical `BASE TABLE` rows for that same identifier. The connector log emits the same six `including 'highgo.testdb.test_a'` entries, so the duplicate is already present at table discovery. Current `dev` at `5f7f5c1dbbb490436605cef37b47a5b30d5acdd1` still has `PostgresDialect.discoverDataCollections()` call `TableDiscoveryUtils.listTables()`. That helper appends every included `INFORMATION_SCHEMA.TABLES` row, and the later map construction then rejects the repeated `TableId`. This is therefore a confirmed discovery bug for HighGo, not a duplicate physical relation and not something to mask with a merge function at the downstream map. I did not find an open PR changing this discovery path. A focused fix should normalize exact duplicate `TableId` values at the discovery boundary while preserving the catalog/schema/table identity; it should not change the downstream map to silently select a winner. Please extend the existing PostgreSQL CDC test coverage with a duplicate-row regression and confirm that an ordinary PostgreSQL discovery result is unchanged. A HighGo-backed integration case would be valuable if its JDBC driver can be made available in the test environment. The issue remains open for that focused fix. Please link a PR here before starting work so that we keep a single implementation path. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
