andygrove opened a new issue, #6259: URL: https://github.com/apache/datafusion-comet/issues/6259
### Describe the bug Two settings accept values that fail only once a task runs on an executor. `spark.comet.shuffle.jvm.batchSize` is only checked to be no larger than `spark.comet.batchSize` ([CometConf.scala#L677-L688](https://github.com/apache/datafusion-comet/blob/bc4be39964cbe9cdb5f2a949740a8164e6b5755b/spark/src/main/scala/org/apache/comet/CometConf.scala#L677-L688)), so 0 is accepted. `process_sorted_row_partition` then never advances, because `n = min(batch_size, ...)` is 0 on every pass of its loop ([row.rs#L1395-L1396](https://github.com/apache/datafusion-comet/blob/bc4be39964cbe9cdb5f2a949740a8164e6b5755b/native/shuffle/src/spark_unsafe/row.rs#L1395-L1396)). The loop runs inside a JNI call, so killing the task doesn't stop it. `spark.comet.exec.memoryPool` has no `checkValues` ([CometConf.scala#L905-L913](https://github.com/apache/datafusion-comet/blob/bc4be39964cbe9cdb5f2a949740a8164e6b5755b/spark/src/main/scala/org/apache/comet/CometConf.scala#L905-L913)). A misspelled value, including a difference in case only, gets through, and in off-heap mode every task then fails with `Unsupported memory pool type` when it creates a native plan. ### Expected behavior Both are validated when they are set. The batch size must be positive, and the pool type must be `fair_unified` or `greedy_unified`, accepted in any case and lowercased before it reaches native code. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
