deniskuzZ commented on PR #6794:
URL: https://github.com/apache/hive/pull/6794#issuecomment-5888545625

   I think the marker column is the wrong layer for this: the information it 
adds is already in the plan, in the per-file delete marker that the t CTE 
produces
   
   COWWithClauseBuilder already groups the matched rows by file. It emits one 
row per affected file (row_number() over (partition by file_path) ... where 
rn=1) with ROW__POSITION hard-coded to -1. The writer only checks pos < 0. So 
the marker can carry the per-file matched count at no extra cost
   
   ````
   t AS (
     select <acid cols, ROW__POSITION := -cnt>, ...
     from (
       select ..., row_number() over (partition by file_path) rn,
                   count(*)     over (partition by file_path) cnt
       from target where <cond>
     ) q
     where rn = 1
   )
   ````
   
   and
   ````
   if (positionDelete.pos() < 0) {
     matchedRows += -positionDelete.pos();   // rows this operation replaced in 
that file
     ...
     replacedDataFiles.add(dataFile);
   }
   ````


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to