DanielLeens commented on issue #12267:
URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5749724877

   Thanks for supplying the complete catalog result. This confirms the 
diagnosis.
   
   There is one physical `testdb.test_a` relation in `pg_class`/`pg_namespace` 
(OID `334287`), while HighGo's `information_schema.tables` returns six 
identical `BASE TABLE` rows for that same identifier. The connector log emits 
the same six `including 'highgo.testdb.test_a'` entries, so the duplicate is 
already present at table discovery.
   
   Current `dev` at `5f7f5c1dbbb490436605cef37b47a5b30d5acdd1` still has 
`PostgresDialect.discoverDataCollections()` call 
`TableDiscoveryUtils.listTables()`. That helper appends every included 
`INFORMATION_SCHEMA.TABLES` row, and the later map construction then rejects 
the repeated `TableId`. This is therefore a confirmed discovery bug for HighGo, 
not a duplicate physical relation and not something to mask with a merge 
function at the downstream map.
   
   I did not find an open PR changing this discovery path. A focused fix should 
normalize exact duplicate `TableId` values at the discovery boundary while 
preserving the catalog/schema/table identity; it should not change the 
downstream map to silently select a winner. Please extend the existing 
PostgreSQL CDC test coverage with a duplicate-row regression and confirm that 
an ordinary PostgreSQL discovery result is unchanged. A HighGo-backed 
integration case would be valuable if its JDBC driver can be made available in 
the test environment.
   
   The issue remains open for that focused fix. Please link a PR here before 
starting work so that we keep a single implementation path.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to