deniskuzZ commented on PR #6794:
URL: https://github.com/apache/hive/pull/6794#issuecomment-5888545625
I think the marker column is the wrong layer for this: the information it
adds is already in the plan, in the per-file delete marker that the t CTE
produces
COWWithClauseBuilder already groups the matched rows by file. It emits one
row per affected file (row_number() over (partition by file_path) ... where
rn=1) with ROW__POSITION hard-coded to -1. The writer only checks pos < 0. So
the marker can carry the per-file matched count at no extra cost
````
t AS (
select <acid cols, ROW__POSITION := -cnt>, ...
from (
select ..., row_number() over (partition by file_path) rn,
count(*) over (partition by file_path) cnt
from target where <cond>
) q
where rn = 1
)
````
and
````
if (positionDelete.pos() < 0) {
matchedRows += -positionDelete.pos(); // rows this operation replaced in
that file
...
replacedDataFiles.add(dataFile);
}
````
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]