andygrove commented on PR #5634: URL: https://github.com/apache/datafusion-comet/pull/5634#issuecomment-5998507917
#6454 is fixed by #6459, which is now merged into this branch, and the 3.5, 4.0, 4.1 and 4.2 diffs no longer skip Spark's `SPARK-42101: Coalesce shuffle partition with union even if exists TableCacheQueryStage`. Its port in `CometInMemoryCacheSuite` passes on Spark 3.5 and 4.1. #6459 also addresses its own review. It hands Spark's `CoalesceShufflePartitions` the whole stage, so a Cartesian product or nested loop join above the union keeps Spark's smaller target size there. It also leaves alone a union whose partitioning an operator above can rely on: on Spark 4.1, coalescing one branch of such a union made Comet's aggregate return each key twice. Main is merged in again too. Since #6415, `BroadcastJoinSuite` runs Comet in the session it builds itself, so its `SPARK-23214` and `SPARK-37742` tests now also count Comet's cache scan and broadcast join. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
