DanielLeens commented on issue #12267:
URL: https://github.com/apache/seatunnel/issues/12267#issuecomment-5731071873

   Thanks for providing the plain-text catalog results. They rule out a 
duplicate row for `test_a` in the individual `pg_class` and `pg_namespace` 
queries, so the immediate fault is now narrower.
   
   Current `dev` at `1325a44b2f4013198fdb2c12f9df345de1cde4bc` still follows 
this path: `PostgresDialect.discoverDataCollections()` calls 
`TableDiscoveryUtils.listTables()`, which first reads `pg_database` and then 
queries each database's `INFORMATION_SCHEMA.TABLES`; 
`PostgresIncrementalSource.tableChanges()` then puts that discovered `TableId` 
list into `Collectors.toMap` without a merge function. The exception therefore 
means the duplicate has entered the connector's discovered-table list before 
schema serialization. A generic `distinct()` or map merge would be unsafe here 
because it could hide an incorrect cross-catalog discovery result.
   
   Please provide these sanitized observations from the failing run:
   
   1. the connector log line listing available databases and every `including 
'...' for further processing` line;
   2. `SELECT datname FROM pg_database ORDER BY datname;`; and
   3. for each catalog listed by the connector, the rows returned for 
`testdb.test_a` from that catalog's `INFORMATION_SCHEMA.TABLES`, including 
`table_catalog`, `table_schema`, `table_name`, and `table_type`.
   
   That will show whether HighGo exposes the same relation through more than 
one catalog query or whether the duplicate is introduced by the connector's 
filtering path. Please also rotate any credentials from the earlier 
configuration if they were real, and keep using redacted values in further 
updates. Once the discovered-list source is proven, the regression should cover 
both an identical duplicate (deterministic handling) and conflicting metadata 
(a descriptive failure rather than silent selection).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to