SEZ9 commented on issue #12267: URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5754889349
Thanks @1362227089, the combined `pg_class`/`pg_namespace` result settles it. There is a single relation for `testdb.test_a` (OID `334287`), while on the `highgo` database `information_schema.tables` returns six identical `BASE TABLE` rows for the same identifier, and the connector log shows the same six `including 'highgo.testdb.test_a' for further processing` lines. So the duplicate is introduced at table discovery by HighGo's `information_schema` view, not by a duplicated table at the storage layer. The empty result on `cflag` is consistent with that as well. On `dev` at `5f7f5c1dbbb490436605cef37b47a5b30d5acdd1`, `PostgresDialect.discoverDataCollections()` goes through `TableDiscoveryUtils.listTables()`, which appends every included `INFORMATION_SCHEMA.TABLES` row, and the later map construction rejects the repeated `TableId`. The fix should drop exact duplicate `TableId` values at that discovery boundary (keeping catalog/schema/table identity intact) rather than teaching the downstream map to pick a winner. Would you like to open the PR for this? If so, please link it here first so we keep a single implementation path. It should include a regression test in the existing PostgreSQL CDC coverage that feeds duplicate discovery rows and asserts a single `TableId`, plus confirmation that plain PostgreSQL discovery output is unchanged. A HighGo-backed integration case would also be welcome if the HighGo JDBC driver can be made available in the test environment — please let us know whether that is feasible on your side. Leaving the issue open for that focused fix. <!-- streview-comment:1209 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
