hudi-agent commented on code in PR #19853:
URL: https://github.com/apache/hudi/pull/19853#discussion_r4000025075


##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/HoodieProcedureFilterUtils.scala:
##########
@@ -411,8 +579,17 @@ object HoodieProcedureFilterUtils {
       // Spark raises SparkArithmeticException for an overflowing ANSI cast or 
arithmetic, and
       // SparkNumberFormatException or SparkDateTimeException for an ANSI cast 
of a malformed
       // string; each extends the matching JDK type. Swallowing one would 
silently drop a row the
-      // same query keeps, so let it out and let the caller fail the way the 
equivalent query does.
+      // same query keeps, so let it out unconditionally, exactly as before 
the registry fallback.
       case Failure(e @ (_: ArithmeticException | _: NumberFormatException | _: 
DateTimeException)) => throw e
+      // SparkThrowable covers the equivalent runtime errors from 
registry-resolved functions
+      // (to_number/bit_get out-of-range, ...), and IllegalArgumentException 
covers a bad regex
+      // pattern - both newly reachable through the registry fallback, so this 
guard only applies
+      // to these two: a caller that skips validateFilterExpression and 
evaluates a genuinely
+      // unsupported function directly still no-matches instead of hitting the 
same
+      // SparkThrowable-family INTERNAL_ERROR that Unevaluable.eval() raises 
for an unrelated
+      // reason, so this method stays safe to call on its own.
+      case Failure(e @ (_: SparkThrowable | _: IllegalArgumentException)) if 
!boundExpr.exists(_.isInstanceOf[Unevaluable]) =>

Review Comment:
   🤖 Could this guard now surface a raw INTERNAL_ERROR for `name ILIKE 'a%'`? 
The parser emits `ILike` directly (it's `RuntimeReplaceable` on 3.3 through 
4.2, never an `UnresolvedFunction`), so `unwrapReplacements` never sees it; it 
passes `validateFilterExpression` (resolved, not `Unevaluable`), and then 
`RuntimeReplaceable.eval` throws a `SparkThrowable` that this arm rethrows - 
where before the PR the rows were just silently dropped. It might be worth 
running the RuntimeReplaceable unwrap over the whole tree (e.g. as part of pass 
three) so ILIKE actually evaluates via its `Like(Lower(l), Lower(r))` 
replacement, or at least having validation reject any leftover 
`RuntimeReplaceable` with a clean message.
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



##########
hudi-spark-datasource/hudi-spark/src/main/scala/org/apache/spark/sql/hudi/command/procedures/HoodieProcedureFilterUtils.scala:
##########
@@ -389,6 +416,147 @@ object HoodieProcedureFilterUtils {
     }
   }
 
+  // Resolves a function not covered by the hardcoded table above via Spark's 
own FunctionRegistry,
+  // then checks the result is actually usable outside a real query plan - 
both steps a plain
+  // lookupFunction call skips or can't tell on its own. Anything that isn't 
falls through to the
+  // existing rejection path instead of letting eval() throw silently.
+  private def resolveViaFunctionRegistry(unresolvedFunc: UnresolvedFunction, 
sparkSession: SparkSession): Expression = {
+    Try {
+      val castedResolved = 
applySparkAnalyzerCoercionRules(lookupBuiltin(unresolvedFunc, sparkSession))
+      // Checked here, on the raw wrapper, before unwrapping: a 
RuntimeReplaceable wrapper's own
+      // declared input-type contract (nvl needing matching operand types, 
split_part needing
+      // string/string/int) is otherwise discarded once unwrapped to a form 
with a weaker or
+      // absent contract of its own.
+      if (!castedResolved.checkInputDataTypes().isSuccess) {
+        unresolvedFunc
+      } else {
+        val finalized = finalizeRegistryResolution(castedResolved)
+        if (isUsableOutsideQueryPlan(finalized)) finalized else unresolvedFunc
+      }
+    }.getOrElse(unresolvedFunc)
+  }
+
+  // Runs a handful of the analyzer's own coercion rules on a single 
expression, the same rules
+  // lookupFunction skips: ImplicitTypeCasts for nodes declaring a real 
input-type contract,
+  // FunctionArgumentConversion/ConcatCoercion/IfCoercion for the builtins 
whose argument types get
+  // unified before the type check even sees them (concat(id, 'x') casts the 
Int to String via
+  // ConcatCoercion in a real query; without it Concat.checkInputDataTypes 
just fails). Order
+  // mirrors TypeCoercion's own rule list - ImplicitTypeCasts last, as a 
catch-all. Anything not
+  // covered by one of these four passes through unchanged.
+  private def applySparkAnalyzerCoercionRules(expression: Expression): 
Expression = {
+    val engine: TypeCoercionBase = if (SQLConf.get.ansiEnabled) 
AnsiTypeCoercion else TypeCoercion
+    Seq(engine.FunctionArgumentConversion, engine.ConcatCoercion, 
engine.IfCoercion, engine.ImplicitTypeCasts)
+      .foldLeft(expression) { (expr, rule) => rule.transform.applyOrElse(expr, 
identity[Expression]) }
+  }
+
+  // Filter expressions only ever call plain builtins. A db-qualified or 3+ 
part name (db.func,
+  // catalog.db.func) can only be resolved by guessing which part is the real 
function name - that
+  // risks matching an unrelated same-named function, so those are left 
unresolved instead.
+  //
+  // Each argument gets widened before the lookup, not just the call's own 
result afterward: an
+  // argument like ts + 1 (Long + Int) is still an unresolved Add at this 
point, and the wrapper's
+  // checkInputDataTypes right after runs before pass three ever gets a chance 
to widen it -
+  // sqrt(ts + 1) would fail that check for the same reason nvl(ts, 0) needed 
pre-lookup widening,
+  // while the hardcoded abs(ts + 1) already works because pass three widens 
its argument too, just
+  // later in the pipeline.
+  private def lookupBuiltin(unresolvedFunc: UnresolvedFunction, sparkSession: 
SparkSession): Expression =
+    unresolvedFunc.nameParts match {
+      case Seq(funcName) =>
+        val widenedArguments = 
unresolvedFunc.arguments.map(applyHudiWideningRules)
+        sparkSession.sessionState.functionRegistry
+          .lookupFunction(builtinFunctionIdentifier(funcName), 
widenedArguments)
+      case _ => unresolvedFunc
+    }
+
+  // A session's function registry is a clone of FunctionRegistry.builtin, 
keyed the same way
+  // builtins are actually registered - a bare name pre-4.2, but the fully 
qualified
+  // system.builtin.<name> from 4.2 onward, where a session-level clone 
(unlike the builtin
+  // singleton itself) stops auto-qualifying a bare name it's given and 
asserts instead.
+  private def builtinFunctionIdentifier(funcName: String): FunctionIdentifier =
+    if (HoodieSparkUtils.gteqSpark4_2) {
+      // FunctionIdentifier only gained the catalog parameter from Spark 3.4 
onward - this file
+      // still compiles against 3.3 too, where the case class has just 
funcName/database, so a
+      // direct 3-arg call wouldn't compile there. Reached through reflection 
instead, the same way
+      // the With handling below reaches classes that don't exist on every 
targeted version.
+      classOf[FunctionIdentifier]
+        .getConstructor(classOf[String], classOf[Option[_]], 
classOf[Option[_]])
+        .newInstance(funcName, Some("builtin"), Some("system"))
+        .asInstanceOf[FunctionIdentifier]
+    } else {
+      FunctionIdentifier(funcName)
+    }
+
+  // RuntimeReplaceable placeholders (nvl, ifnull, left, right, ...) need 
substitution the analyzer
+  // normally performs but lookupFunction skips, and can themselves unwrap to 
another
+  // RuntimeReplaceable (regexp_substr -> NullIf), so the unwrap runs to a 
fixed point. Then widens
+  // numeric operands the same way pass three would - nvl(ts, 0) unwraps to 
Coalesce(ts, 0), which
+  // needs the same widening the hardcoded coalesce(ts, 0) case gets - so a 
registry function and
+  // its hardcoded-table equivalent agree on what counts as resolved.
+  //
+  // A replacement can itself be a With(child, defs) common-subexpression 
wrapper from 4.0 onward
+  // (NullIf's, for instance) - Unevaluable like any other holder, so it needs 
inlining here too or
+  // isUsableOutsideQueryPlan rejects it outright. 
With/CommonExpressionDef/CommonExpressionRef
+  // don't exist before 4.0, so this goes by reflection rather than a direct 
import; evaluating a
+  // filter once per row rather than once per query makes the dedup With 
exists for irrelevant, so
+  // inlining each reference in place of its definition is exactly equivalent 
to keeping it.
+  private def finalizeRegistryResolution(expression: Expression): Expression = 
{
+    def unwrapReplacements(expr: Expression): Expression = {
+      val next = expr.transformUp {
+        case r: RuntimeReplaceable => r.replacement
+        case withExpr if isWithNode(withExpr) => 
inlineCommonExpressions(withExpr)
+      }
+      if (next.fastEquals(expr)) next else unwrapReplacements(next)
+    }
+    applyHudiWideningRules(unwrapReplacements(expression))
+  }
+
+  // Reflection here for the same reason finalizeRegistryResolution's comment 
above gives.
+  private def isWithNode(expression: Expression): Boolean = 
expression.getClass.getSimpleName == "With"

Review Comment:
   🤖 nit: could `isWithNode` / `isCommonExpressionRef` compare against the 
fully qualified class name (e.g. 
`"org.apache.spark.sql.catalyst.expressions.With"`) instead of `getSimpleName`? 
A bare `"With"` could match an unrelated class, and a shared 
`isCatalystClass(expr, fqcn)` helper would keep the two checks in one place.
   
   <sub><i>⚠️ AI-generated; verify before applying. React 👍/👎 to flag 
quality.</i></sub>



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to