[
https://issues.apache.org/jira/browse/HUDI-4946?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18013628#comment-18013628
]
Danny Chen commented on HUDI-4946:
----------------------------------
> I think merge into for the same table need has consistent behavior.
This is corrent but the fix does not make sense, it does not duplicate when
there is no matched row, MIT should always use `UPSERT` operation, the incoming
data set should always be dedupped, that is how to keep consistent with MIT
that has matched rows.
> merge into with no preCombineField has dup row in only insert
> -------------------------------------------------------------
>
> Key: HUDI-4946
> URL: https://issues.apache.org/jira/browse/HUDI-4946
> Project: Apache Hudi
> Issue Type: Bug
> Components: spark-sql
> Reporter: Knight Chess
> Assignee: Knight Chess
> Priority: Minor
> Labels: pull-request-available
> Fix For: 0.12.2
>
>
> a table with precombineField use diff merge into sql has different result.
> If merge sql with matchAction will have no dup key result, but when merge sql
> only has insertAction will have dup result if we not set
> `hoodie.combine.before.insert=true`.
> I think merge into for the same table need has consistent behavior.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)