[PR] [SPARK-48065][SQL] SPJ: allowJoinKeysSubsetOfPartitionKeys is too strict [spark]

via GitHub Wed, 01 May 2024 13:51:54 -0700


szehon-ho opened a new pull request, #46325:
URL: https://github.com/apache/spark/pull/46325


     ### What changes were proposed in this pull request?
   If spark.sql.v2.bucketing.allowJoinKeysSubsetOfPartitionKeys.enabled is 
true, change KeyGroupedPartitioning.satisfies0(distribution) check from all 
clustering keys (here, join keys)  being in partition keys, to the two sets 
overlapping.
   
     ### Why are the changes needed?
   If spark.sql.v2.bucketing.allowJoinKeysSubsetOfPartitionKeys.enabled is 
true, then SPJ no longer triggers if there are more join keys than partition 
keys. But SPJ is supported in this case if flag is false.
   
     ### Does this PR introduce _any_ user-facing change?
   No
   
     ### How was this patch tested?
   Added tests in KeyGroupedPartitioningSuite
   
    ### Was this patch authored or co-authored using generative AI tooling?
   No
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: reviews-unsubscr...@spark.apache.org

For queries about this service, please contact Infrastructure at:
us...@infra.apache.org


---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscr...@spark.apache.org
For additional commands, e-mail: reviews-h...@spark.apache.org

[PR] [SPARK-48065][SQL] SPJ: allowJoinKeysSubsetOfPartitionKeys is too strict [spark]

Reply via email to