adriangb opened a new pull request, #26142:
URL: https://github.com/apache/datafusion/pull/26142

   ## Which issue does this PR close?
   
   - Closes #26141.
   
   ## Rationale for this change
   
   `RepartitionExec` does not apply `datafusion.execution.spill_compression`, 
so its spill files are always uncompressed. All other spilling operators (sort, 
aggregate, sort-merge join, nested loop join) apply the setting. A user who 
sets a codec to decrease spill disk usage gets no decrease for the data that 
`RepartitionExec` spills.
   
   ## What changes are included in this PR?
   
   - `RepartitionExec::execute` now builds its `SpillManager` with 
`.with_compression_type(context.session_config().spill_compression())`, the 
same as the other operators.
   - No change is necessary on the read side. Spill files use the Arrow IPC 
Stream format, and the reader gets the codec from the stream.
   - I checked the other non-test `SpillManager::new` call sites. 
`RepartitionExec` was the only one that did not pass the setting.
   
   ## What is the testing strategy for this PR?
   
   The new test `repartition_spill_honors_spill_compression` runs a spilling 
`RepartitionExec` on constant (highly compressible) data with `uncompressed`, 
`lz4_frame` and `zstd`. It asserts that the same rows spill, that the data 
reads back intact, and that `spilled_bytes` with a codec is less than half of 
the uncompressed value.
   
   The test fails on `main` (`679264 vs 679264 bytes`) and passes with the fix:
   
   | `spill_compression` | `spilled_bytes` on `main` | `spilled_bytes` with 
this PR |
   | --- | --- | --- |
   | `uncompressed` | 679,264 | 679,264 |
   | `lz4_frame` | 679,264 | 7,744 |
   | `zstd` | 679,264 | 5,024 |
   
   ## Are there any user-facing changes?
   
   Yes. Spill files from `RepartitionExec` now use the codec from 
`datafusion.execution.spill_compression`. The default (`uncompressed`) does not 
change. There are no API changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to