pedrumj2 commented on issue #13014:
URL: https://github.com/apache/gluten/issues/13014#issuecomment-5670711527

   Hi @zhouyuan  thanks for the quick response and sharing the context around 
the migration use case. 
   
   > `CREATE TEMPORARY FUNCTION ` IIRC this is a must otherwise Spark will 
raise exceptions unless we bypass this in Spark analyzer?
   
   I created a PR demonstrating the changes needed to make this work without 
bypassing the analyzer, please see 
[PR13016](https://github.com/apache/gluten/pull/13016). I verified by running 
the query e2e (please see the test section of the PR). Spark's 
[SparkSessionExtensions.injectFunction](https://github.com/apache/spark/blob/v3.5.5/sql/core/src/main/scala/org/apache/spark/sql/SparkSessionExtensions.scala#L365)
 exposes a public API for adding entries to a session's FunctionRegistry, and 
the PR uses this API to register native UDF names through it. Appears this was 
a working feature until it was taken out in 
[PR7016](https://github.com/apache/gluten/pull/7016/changes).
   
   > with the new design how do we distinguish same function with different 
parameters, e.g. my_udf(string) vs my_udf(bigint) 
   
   Good question, please see my 
[PR13016](https://github.com/apache/gluten/pull/13016). It calls the  
[getUdfExpression](https://github.com/apache/gluten/blob/9f6dcb599d3ebb8ef9f55450163aae4f025300e9/backends-bolt/src/main/scala/org/apache/spark/sql/expression/UDFResolver.scala#L341)
 which is able to handle the UDF overload. Also tested locally and verified UDF 
overloads work
   
   ```
     (children: Seq[Expression]) => getUdfExpression(name, name)(children))
   ```
   
   This is the output of the local test but I can share the full script as well:
   ```
   $ cat "$WORK/overload.sql"
   SELECT
     col1,
     col2,
     myudf_overload(col1)      AS inc_bigint,
     myudf_overload(col2)      AS inc_varchar,
     myudf_overload(col1, 10L) AS add_bigint_bigint
   FROM VALUES (1L, 'a'), (2L, 'b'), (3L, 'c') AS t(col1, col2)
   ORDER BY col1;
   EXPLAIN SELECT
     myudf_overload(col1)      AS inc_bigint,
     myudf_overload(col2)      AS inc_varchar,
     myudf_overload(col1, 10L) AS add_bigint_bigint
   FROM VALUES (1L, 'a') AS t(col1, col2);
   ### Step 8 — run, verbatim from the walkthrough
   exit status 0
   26/09/14 20:05:10 WARN NativeCodeLoader: Unable to load native-hadoop 
library for your platform... using builtin-java classes where applicable
   Setting default log level to "WARN".
   To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use 
setLogLevel(newLevel).
   W20260914 20:05:12.204257   161 MemoryArbitrator.cpp:84] [MEM] Query memory 
capacity[921.00MB] is set for NOOP arbitrator which has no capacity enforcement
   26/09/14 20:05:13 WARN HiveConf: HiveConf of name hive.stats.jdbc.timeout 
does not exist
   26/09/14 20:05:13 WARN HiveConf: HiveConf of name hive.stats.retries.wait 
does not exist
   26/09/14 20:05:15 WARN ObjectStore: Version information not found in 
metastore. hive.metastore.schema.verification is not enabled so recording the 
schema version 2.3.0
   26/09/14 20:05:15 WARN ObjectStore: setMetaStoreSchemaVersion called but 
recording version is disabled: version = 2.3.0, comment = Set by MetaStore 
[email protected]
   26/09/14 20:05:15 WARN ObjectStore: Failed to get database default, 
returning NoSuchObjectException
   Spark master: local[2], Application Id: local-1789416312258
   26/09/14 20:05:16 WARN SparkShimProvider: Spark runtime version 3.5.9 is not 
matched with Gluten's fully tested version 3.5.5
   26/09/14 20:05:17 WARN GlutenFallbackReporter: Validation failed for plan: 
LocalTableScan[QueryId=0], due to: Gluten does not touch it or does not support 
it
   26/09/14 20:05:19 WARN GlutenFallbackReporter: Validation failed for plan: 
LocalTableScan[QueryId=0], due to: Gluten does not touch it or does not support 
it
   1    a       2       a+1     11
   2    b       3       b+1     12
   3    c       4       c+1     13
   Time taken: 3.316 seconds, Fetched 3 row(s)
   26/09/14 20:05:19 WARN GlutenFallbackReporter: Validation failed for plan: 
LocalTableScan[QueryId=1], due to: Gluten does not touch it or does not support 
it
   == Physical Plan ==
   VeloxColumnarToRow
   +- ^(1) ProjectExecTransformer [myudf_overload(myudf_overload, 
myudf_overload, LongType, true, col1#18L) AS inc_bigint#10L, 
myudf_overload(myudf_overload, myudf_overload, StringType, true, col2#19) AS 
inc_varchar#11, myudf_overload(myudf_overload, myudf_overload, LongType, true, 
col1#18L, 10) AS add_bigint_bigint#12L]
      +- ^(1) InputIteratorTransformer[col1#18L, col2#19]
         +- RowToVeloxColumnar
            +- LocalTableScan [col1#18L, col2#19]
   ```
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to