ryukobayashi commented on PR #6794:
URL: https://github.com/apache/hive/pull/6794#issuecomment-5906627101

   @deniskuzZ Thanks for the detailed review. I agree that adding a trailing 
marker column is not the right layer.
   
   I changed the implementation to use the existing negative ROW__POSITION 
marker instead. COWWithClauseBuilder now calculates count(*) over the file_path 
partition and emits one replacement marker per file with ROW__POSITION = 
-matched_count. HiveIcebergCopyOnWriteRecordWriter sums the absolute values of 
these negative positions and exposes the result through a dedicated 
affected-row writer capability.
   
   FileSinkOperator uses the affected-row count only when all output writers 
support the capability. Otherwise, it falls back to the existing physical 
output-row count, so regular INSERTs and non-CoW writers retain their previous 
behavior. This also avoids adding a column to the query projection or passing 
an extra field through the shuffle and SerDe.
   
   For MERGE, the affected-row count follows the semantics described in the 
review: it counts matched target rows, i.e. UPDATE + DELETE rows, and does not 
include INSERT rows. Therefore, the test case with two updates, one delete, and 
two inserts now expects 3 affected rows.
   
   The COW UPDATE and MERGE regression tests pass, and the affected Iceberg 
qtests were regenerated and pass with the new execution plans.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to