Fang-Yu Rao has posted comments on this change. ( http://gerrit.cloudera.org:8080/24137 )
Change subject: IMPALA-13810: Generate lineage records for UPDATE and MERGE statements ...................................................................... Patch Set 5: (3 comments) I left some comments on the added test cases. Let me know if I missed something important. Thanks! http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test File testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test: http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@157 PS5, Line 157: "operationType": "UPDATE" For the UPDATE queries, conceptually the source and destination tables correspond to the same table. But it looks like we have one source table and one destination table as if they were 2 different tables. Is there a reason why we need 2 tables for constructing a lineage graph in the case of UPDATE? http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@582 PS5, Line 582: "queryText": "merge into lineage_dml_test_db.iceberg_non_partitioned t using functional_parquet.iceberg_non_partitioned s on t.id = s.id when matched then update set t.user = s.user when not matched then insert values (s.id, s.user, s.action, s.event_time)", In this query, we have one source table, and one destination table. Thus a table has 4 columns, each corresponding to one vertex. But in the respective lineage graph, we have 11 vertices in total. I took a closer look, and found that some vertices like vertex 0 and vertex 2 correspond to the same column 'id' (of the same table 'lineage_dml_test_db.iceberg_non_partitioned'). Is this expected? http://gerrit.cloudera.org:8080/#/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test@1448 PS5, Line 1448: "queryText": "update lineage_dml_test_db.kudu_tbl set string_col = \"foo\" where float_col < 0.5 and bool_col = true", : "operationType": "UPDATE" It looks like for the UPDATE queries against Kudu tables, we do not list all the columns in the tables as vertices. Take this query for example. We have 14 columns (including the column 'auto_incrementing_id'). It was mentioned in the commit message that we do not include the auto incrementing id column in the lineage graph, but I am wondering why some columns are not listed as vertices, e.g., tinyint_col. Is that because the column 'tinyint_col' is not explicitly referenced in the query? But if this is the case, why do we list all columns in the table 'lineage_dml_test_db.iceberg_non_partitioned' as vertices even though the column 'event_time' was not explicitly referenced in the query at https://gerrit.cloudera.org/c/24137/5/testdata/workloads/functional-query/queries/QueryTest/lineage-update-merge.test#156. -- To view, visit http://gerrit.cloudera.org:8080/24137 To unsubscribe, visit http://gerrit.cloudera.org:8080/settings Gerrit-Project: Impala-ASF Gerrit-Branch: master Gerrit-MessageType: comment Gerrit-Change-Id: Icba79f756509438455bfbf3067733f6f29284220 Gerrit-Change-Number: 24137 Gerrit-PatchSet: 5 Gerrit-Owner: Daniel Vanko <[email protected]> Gerrit-Reviewer: Daniel Vanko <[email protected]> Gerrit-Reviewer: Fang-Yu Rao <[email protected]> Gerrit-Reviewer: Impala Public Jenkins <[email protected]> Gerrit-Reviewer: Peter Rozsa <[email protected]> Gerrit-Reviewer: Zoltan Borok-Nagy <[email protected]> Gerrit-Comment-Date: Tue, 05 May 2026 00:40:19 +0000 Gerrit-HasComments: Yes
