acvictor opened a new pull request, #13072: URL: https://github.com/apache/gluten/pull/13072
## What changes are proposed in this pull request? ColumnarShuffleManager built its IndexShuffleBlockResolver with the single-arg constructor, so the resolver allocated its own taskIdMapsForShuffle while the manager kept a second, separate map. The resolver records blocks migrated in during executor decommissioning into its map, but unregisterShuffle reads the manager's map to delete map output. With two maps, migrated blocks were recorded where nothing read them, so their files were never deleted and leaked disk on decommissioned executors. Spark wires a single shared map in SortShuffleManager. taskIdMapsForShuffle is now declared before shuffleBlockResolver -- order matters, since Scala initializes vals in declaration order and the previous ordering would capture null -- and passed to the resolver. unregisterShuffle now iterates under mapTaskIds.synchronized, matching Spark, because the block-migration path mutates the same set under that lock. stop() now defers to super.stop(), which stops the resolver. Adds a ColumnarShuffleManagerSuite case that inserts an entry through the resolver's map and asserts unregisterShuffle clears it, which fails if the two maps are ever unshared again. ## How was this patch tested? UT ## Was this patch authored or co-authored using generative AI tooling? Co-authored by GitHub Copilot CLI 1.0.86 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
