yabola opened a new issue, #6830:
URL: https://github.com/apache/kyuubi/issues/6830

   ### Code of Conduct
   
   - [X] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   
   
   ### Search before asking
   
   - [X] I have searched in the 
[issues](https://github.com/apache/kyuubi/issues?q=is%3Aissue) and found no 
similar issues.
   
   
   ### What would you like to be improved?
   
   when merging small files(set 
`spark.sql.optimizer.insertRepartitionBeforeWrite.enabled`=true) , the default 
session advisory partition size (64MB) will be used as target. This default 
value can still lead to small files because the written data can be compressed 
nicely using columnar file formats (usually 1/4 or smaller of the shuffle 
exchange size, the result is often around 15MB).
   
   Spark now support configuring the rebalance expression advisory size in 
https://github.com/apache/spark/pull/40421 . So we can have a configuration 
that can configure the merge size separately.
   
   ### How should we improve?
   
   add one configuration to control the size separately
   
   ### Are you willing to submit PR?
   
   - [X] Yes. I would be willing to submit a PR with guidance from the Kyuubi 
community to improve.
   - [ ] No. I cannot submit a PR at this time.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to