ryukobayashi commented on PR #6794: URL: https://github.com/apache/hive/pull/6794#issuecomment-5906627101
@deniskuzZ Thanks for the detailed review. I agree that adding a trailing marker column is not the right layer. I changed the implementation to use the existing negative ROW__POSITION marker instead. COWWithClauseBuilder now calculates count(*) over the file_path partition and emits one replacement marker per file with ROW__POSITION = -matched_count. HiveIcebergCopyOnWriteRecordWriter sums the absolute values of these negative positions and exposes the result through a dedicated affected-row writer capability. FileSinkOperator uses the affected-row count only when all output writers support the capability. Otherwise, it falls back to the existing physical output-row count, so regular INSERTs and non-CoW writers retain their previous behavior. This also avoids adding a column to the query projection or passing an extra field through the shuffle and SerDe. For MERGE, the affected-row count follows the semantics described in the review: it counts matched target rows, i.e. UPDATE + DELETE rows, and does not include INSERT rows. Therefore, the test case with two updates, one delete, and two inserts now expects 3 affected rows. The COW UPDATE and MERGE regression tests pass, and the affected Iceberg qtests were regenerated and pass with the new execution plans. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
