pedrumj2 commented on issue #13014: URL: https://github.com/apache/gluten/issues/13014#issuecomment-5670711527
Hi @zhouyuan thanks for the quick response and sharing the context around the migration use case. > `CREATE TEMPORARY FUNCTION ` IIRC this is a must otherwise Spark will raise exceptions unless we bypass this in Spark analyzer? I created a PR demonstrating the changes needed to make this work without bypassing the analyzer, please see [PR13016](https://github.com/apache/gluten/pull/13016). I verified by running the query e2e (please see the test section of the PR). Spark's [SparkSessionExtensions.injectFunction](https://github.com/apache/spark/blob/v3.5.5/sql/core/src/main/scala/org/apache/spark/sql/SparkSessionExtensions.scala#L365) exposes a public API for adding entries to a session's FunctionRegistry, and the PR uses this API to register native UDF names through it. Appears this was a working feature until it was taken out in [PR7016](https://github.com/apache/gluten/pull/7016/changes). > with the new design how do we distinguish same function with different parameters, e.g. my_udf(string) vs my_udf(bigint) Good question, please see my [PR13016](https://github.com/apache/gluten/pull/13016). It calls the [getUdfExpression](https://github.com/apache/gluten/blob/9f6dcb599d3ebb8ef9f55450163aae4f025300e9/backends-bolt/src/main/scala/org/apache/spark/sql/expression/UDFResolver.scala#L341) which is able to handle the UDF overload. Also tested locally and verified UDF overloads work ``` (children: Seq[Expression]) => getUdfExpression(name, name)(children)) ``` This is the output of the local test but I can share the full script as well: ``` $ cat "$WORK/overload.sql" SELECT col1, col2, myudf_overload(col1) AS inc_bigint, myudf_overload(col2) AS inc_varchar, myudf_overload(col1, 10L) AS add_bigint_bigint FROM VALUES (1L, 'a'), (2L, 'b'), (3L, 'c') AS t(col1, col2) ORDER BY col1; EXPLAIN SELECT myudf_overload(col1) AS inc_bigint, myudf_overload(col2) AS inc_varchar, myudf_overload(col1, 10L) AS add_bigint_bigint FROM VALUES (1L, 'a') AS t(col1, col2); ### Step 8 — run, verbatim from the walkthrough exit status 0 26/09/14 20:05:10 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable Setting default log level to "WARN". To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel). W20260914 20:05:12.204257 161 MemoryArbitrator.cpp:84] [MEM] Query memory capacity[921.00MB] is set for NOOP arbitrator which has no capacity enforcement 26/09/14 20:05:13 WARN HiveConf: HiveConf of name hive.stats.jdbc.timeout does not exist 26/09/14 20:05:13 WARN HiveConf: HiveConf of name hive.stats.retries.wait does not exist 26/09/14 20:05:15 WARN ObjectStore: Version information not found in metastore. hive.metastore.schema.verification is not enabled so recording the schema version 2.3.0 26/09/14 20:05:15 WARN ObjectStore: setMetaStoreSchemaVersion called but recording version is disabled: version = 2.3.0, comment = Set by MetaStore [email protected] 26/09/14 20:05:15 WARN ObjectStore: Failed to get database default, returning NoSuchObjectException Spark master: local[2], Application Id: local-1789416312258 26/09/14 20:05:16 WARN SparkShimProvider: Spark runtime version 3.5.9 is not matched with Gluten's fully tested version 3.5.5 26/09/14 20:05:17 WARN GlutenFallbackReporter: Validation failed for plan: LocalTableScan[QueryId=0], due to: Gluten does not touch it or does not support it 26/09/14 20:05:19 WARN GlutenFallbackReporter: Validation failed for plan: LocalTableScan[QueryId=0], due to: Gluten does not touch it or does not support it 1 a 2 a+1 11 2 b 3 b+1 12 3 c 4 c+1 13 Time taken: 3.316 seconds, Fetched 3 row(s) 26/09/14 20:05:19 WARN GlutenFallbackReporter: Validation failed for plan: LocalTableScan[QueryId=1], due to: Gluten does not touch it or does not support it == Physical Plan == VeloxColumnarToRow +- ^(1) ProjectExecTransformer [myudf_overload(myudf_overload, myudf_overload, LongType, true, col1#18L) AS inc_bigint#10L, myudf_overload(myudf_overload, myudf_overload, StringType, true, col2#19) AS inc_varchar#11, myudf_overload(myudf_overload, myudf_overload, LongType, true, col1#18L, 10) AS add_bigint_bigint#12L] +- ^(1) InputIteratorTransformer[col1#18L, col2#19] +- RowToVeloxColumnar +- LocalTableScan [col1#18L, col2#19] ``` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
