Fang-Yu Rao has posted comments on this change. ( 
http://gerrit.cloudera.org:8080/24137 )

Change subject: IMPALA-13810: Generate lineage records for UPDATE and MERGE 
statements
......................................................................


Patch Set 5:

(3 comments)

I left some comments on the added test cases. Let me know if I missed something 
important. Thanks!

http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test
File 
testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test:

http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@157
PS5, Line 157: "operationType": "UPDATE"
For the UPDATE queries, conceptually the source and destination tables 
correspond to the same table. But it looks like we have one source table and 
one destination table as if they were 2 different tables. Is there a reason why 
we need 2 tables for constructing a lineage graph in the case of UPDATE?


http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@582
PS5, Line 582: "queryText": "merge into 
lineage_dml_test_db.iceberg_non_partitioned t using 
functional_parquet.iceberg_non_partitioned s on t.id = s.id when matched then 
update set t.user = s.user when not matched then insert values (s.id, s.user, 
s.action, s.event_time)",
In this query, we have one source table, and one destination table. Thus a 
table has 4 columns, each corresponding to one vertex. But in the respective 
lineage graph, we have 11 vertices in total.

I took a closer look, and found that some vertices like vertex 0 and vertex 2 
correspond to the same column 'id' (of the same table 
'lineage_dml_test_db.iceberg_non_partitioned'). Is this expected?


http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@1448
PS5, Line 1448: "queryText": "update lineage_dml_test_db.kudu_tbl set 
string_col = \"foo\" where float_col < 0.5   and bool_col = true",
              :   "operationType": "UPDATE"
It looks like for the UPDATE queries against Kudu tables, we do not list all 
the columns in the tables as vertices.

Take this query for example. We have 14 columns (including the column 
'auto_incrementing_id'). It was mentioned in the commit message that we do not 
include the auto incrementing id column in the lineage graph, but I am 
wondering why some columns are not listed as vertices, e.g., tinyint_col. Is 
that because the column 'tinyint_col' is not explicitly referenced in the 
query? But if this is the case, why do we list all columns in the table 
'lineage_dml_test_db.iceberg_non_partitioned' as vertices even though the 
column 'event_time' was not explicitly referenced in the query at 
https://gerrit.cloudera.org/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test#156.



--
To view, visit http://gerrit.cloudera.org:8080/24137
To unsubscribe, visit http://gerrit.cloudera.org:8080/settings

Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-MessageType: comment
Gerrit-Change-Id: Icba79f756509438455bfbf3067733f6f29284220
Gerrit-Change-Number: 24137
Gerrit-PatchSet: 5
Gerrit-Owner: Daniel Vanko <[email protected]>
Gerrit-Reviewer: Daniel Vanko <[email protected]>
Gerrit-Reviewer: Fang-Yu Rao <[email protected]>
Gerrit-Reviewer: Impala Public Jenkins <[email protected]>
Gerrit-Reviewer: Peter Rozsa <[email protected]>
Gerrit-Reviewer: Zoltan Borok-Nagy <[email protected]>
Gerrit-Comment-Date: Tue, 05 May 2026 00:40:19 +0000
Gerrit-HasComments: Yes

Reply via email to