pedrumj2 opened a new pull request, #13016:
URL: https://github.com/apache/gluten/pull/13016

   ## What changes are proposed in this pull request?
   
   Resolves #13014. UDFs signatures will directly be read from the C++ velox 
implementation and no longer require registering a separate Java Interface class
   
   ## How was this patch tested?
   
   ### Unit test
   
   `UDFResolverSuite` covers the selection rules: a plain name is described, a 
dotted name and a name colliding with a Spark built-in are not, and the output 
is sorted. It passed locally — `Tests: succeeded 4, failed 0`.
   
   ### Formatting
   
   `./dev/format-scala-code.sh check` passes.
   
   ### End-to-end run
   
   The UDF was also called by its own name. The example library in 
[MyUDF.cc](https://github.com/apache/gluten/blob/main/cpp/velox/udf/examples/MyUDF.cc)
 has no dotless UDF, so one was added locally for the run and is deliberately 
not part of this patch. The steps below mirror those in #13014, so the ones 
this patch removes are marked.
   
   #### [step1] Write the Java UDF
   
   Not needed. This is what the patch removes.
   
   #### [step2] Write the C++ Velox UDF
   
   Added to `cpp/velox/udf/examples/MyUDF.cc`:
   
   ```cpp
   namespace myudfincrement {
   
   template <typename T>
   struct MyUdfIncrementFunction {
     VELOX_DEFINE_FUNCTION_TYPES(T);
   
     FOLLY_ALWAYS_INLINE void call(int64_t& result, const int64_t& a) {
       result = a + 1;
     }
   };
   
   // name: myudf_increment
   // signatures:
   //    bigint -> bigint
   // type: SimpleFunction
   class MyUdfIncrementRegisterer final : public gluten::UdfRegisterer {
    public:
     int getNumUdf() override {
       return 1;
     }
   
     void populateUdfEntries(int& index, gluten::UdfEntry* udfEntries) override 
{
       udfEntries[index++] = {name_.c_str(), kBigInt, 1, arg_, false, false};
     }
   
     void registerSignatures() override {
       facebook::velox::registerFunction<MyUdfIncrementFunction, int64_t, 
int64_t>({name_});
     }
   
    private:
     const std::string name_ = "myudf_increment";
     const char* arg_[1] = {kBigInt};
   };
   
   } // namespace myudfincrement
   ```
   
   registered alongside the existing one:
   
   ```cpp
   
registerers.push_back(std::make_shared<myudfincrement::MyUdfIncrementRegisterer>());
   ```
   
   #### [step3] Build the `.so`
   
   ```bash
   ./dev/builddeps-veloxbe.sh --build_examples=ON
   ```
   
   Produces `cpp/build/velox/udf/examples/libmyudf.so`.
   
   #### [step4] Compile the Java UDF into a jar
   
   Not needed. There is no Java class to compile.
   
   #### [step5] Write the query
   
   No `CREATE TEMPORARY FUNCTION`; the name comes from the loaded library alone.
   
   ```sql
   SELECT col1, myudf_increment(col1) AS incremented
   FROM VALUES (1L), (2L), (3L) AS t(col1)
   ORDER BY col1;
   
   EXPLAIN SELECT myudf_increment(col1) FROM VALUES (1L) AS t(col1);
   ```
   
   #### [step6] Run the query
   
   Only the Gluten bundle is on the classpath.
   
   ```bash
   $SPARK_HOME/bin/spark-sql \
     --master 'local[2]' \
     --jars "$BUNDLE" \
     --driver-class-path "$BUNDLE" \
     --conf spark.plugins=org.apache.gluten.GlutenPlugin \
     --conf spark.memory.offHeap.enabled=true \
     --conf spark.memory.offHeap.size=4g \
     --conf 
spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager \
     --conf 
spark.gluten.sql.columnar.backend.velox.udfLibraryPaths=file:///path/to/libmyudf.so
 \
     --conf 
spark.gluten.sql.columnar.backend.velox.driver.udfLibraryPaths=file:///path/to/libmyudf.so
 \
     -f demo.sql
   ```
   
   #### Result
   
   Spark 3.5.9, with a Gluten bundle jar built from this branch.
   
   ```
   1    2
   2    3
   3    4
   Time taken: 3.326 seconds, Fetched 3 row(s)
   == Physical Plan ==
   VeloxColumnarToRow
   +- ^(1) ProjectExecTransformer [myudf_increment(myudf_increment, 
myudf_increment, LongType, true, col1#9L) AS myudf_increment(col1)#10L]
      +- ^(1) InputIteratorTransformer[col1#9L]
         +- RowToVeloxColumnar
            +- LocalTableScan [col1#9L]
   ```
   
   
[ProjectExecTransformer](https://github.com/apache/gluten/blob/main/gluten-substrait/src/main/scala/org/apache/gluten/execution/ProjectExecTransformer.scala)
 shows the native function was used.
   
   ## Was this patch authored or co-authored using generative AI tooling?
   
   Co authored with Claude Code claude-opus-5
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to