SEZ9 commented on issue #12598: URL: https://github.com/apache/seatunnel/issues/12598#issuecomment-5965126655
Thanks for the detailed report and reproduction. The analysis makes sense: the snapshot chunks are range predicates on the split column, a NULL never matches a range predicate, and there is no chunk that covers NULL values, so rows with NULL in a nullable unique-key split column are skipped once the table is split into more than one chunk. Your proposed direction also sounds reasonable: avoid choosing a nullable column as the split column and fall back to reading the table as a single split when no non-nullable key column is available. As you note, for MySQL the parsed Debezium schema cannot be relied on for nullability here, so the check would need to use the actual database metadata. Thanks for opening #12597 — we can continue the discussion on the implementation details there. <!-- streview-comment:1487 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
