viirya commented on PR #6130: URL: https://github.com/apache/datafusion-comet/pull/6130#issuecomment-5879777305
Thanks for catching both issues. In `67a9f776a`, `CometArrowEvalPythonExec.stringArgs` no longer renders `nativeOp`, so the pickled command and native subtree are absent from the plan string. `equals` and `hashCode` compare the Arrow UDF payload while ignoring the outer `plan_id`. A test checks that the command is absent from the plan, changing only the plan ID preserves equality, and changing the command does not. I also documented that embedded Python and PyArrow allocations are excluded from Rust's `allocated` figure and need additional executor memory overhead. After merging upstream in `de71a8e78`, that guidance lives in the new `tuning/memory.md` page and covers the JVM Arrow metric's treatment of PyArrow buffers imported from native. Preflight and the PyArrow UDF jobs on Spark 4.0, 4.1, and 4.2 pass; the Spark 4.1 SQL gate is still running. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
