sabbasani commented on code in PR #12726: URL: https://github.com/apache/gluten/pull/12726#discussion_r4141805718
########## docs/velox-backend-limitations.md: ########## @@ -21,7 +21,65 @@ Gluten currently doesn't support ANSI mode. If ANSI is enabled, Spark plan's exe We now have a issue tracker on ANSI support progress. Please check [issue-10134](https://github.com/apache/gluten/issues/10134). #### Case Sensitive mode -Gluten only supports spark default case-insensitive mode. If case-sensitive mode is enabled, user may get incorrect result. +Gluten respects Spark's case-sensitive configuration (`spark.sql.caseSensitive`). Since +[GLUTEN-1577](https://github.com/apache/gluten/issues/1577) (merged 2023-05), column-name +normalisation in the core engine uses `ConverterUtils.normalizeColName`, which preserves the +original casing when `caseSensitiveAnalysis=true` and lowercases only when it is `false` (the +Spark default). Standard data operations such as scan, filter, aggregation, and join are +therefore correct in both modes. + +**This change addresses the following identified metadata-name collision paths:** + +- `IcebergScanTransformer`: previously used unconditional `equalsIgnoreCase` in + `getMetadataColumns` and unconditional `toLowerCase` in the read-schema field set, causing a + user data column named `Input_File_Name` (or any mixed-case variant of an Iceberg metadata + column name) to be misclassified as a metadata column under `caseSensitive=true`. + Fixed by switching to `ConverterUtils.normalizeColName` throughout the Iceberg scan path. + Validated by `IcebergSuite` / `VeloxIcebergSuite`. + +- `PushDownInputFileExpression` (core rule, `gluten-substrait`): two unconditional + `toLowerCase` usages — one in `containsInputFileRelatedExpr` and one in the `PostOffload` + deduplication — caused incorrect pre-offload rewriting and dangling-attribute plan errors when + a data column named `Input_File_Name` was projected alongside `input_file_name()` under + `caseSensitive=true`. Fixed by using `ConverterUtils.normalizeColName` for gate detection and + `exprId` identity for deduplication. + Validated by `FallbackSuite` (Velox) and `IcebergSuite`. + +- **Delta optimised writer** (`GlutenDeltaOptimizedWriterExec` / `DeltaOptimizedWriterTransformer`): + previously used `caseInsensitiveResolution` (a hardcoded case-insensitive comparator) for + partition-column lookup, ignoring `spark.sql.caseSensitive=true`. Fixed by switching to + `SQLConf.get.resolver`, which honours the session case-sensitivity setting. + Full end-to-end Delta writer tests require a native Delta backend and are not run in CI for + this module; the resolver-semantics contract is validated at the unit level by + `GlutenClickHouseCaseSensitiveSchemaSuite`. + +- **ClickHouse `CHIteratorApi.getFileSchema`**: previously used `equalsIgnoreCase` for schema + field matching, ignoring `caseSensitive=true`. Fixed by switching to `SQLConf.get.resolver`. + The resolver-semantics contract is validated by `GlutenClickHouseCaseSensitiveSchemaSuite` + (unit-level only; end-to-end requires a running ClickHouse backend). Review Comment: Refactored velox-backend-limitations.md to remove PR-specific change log notes, test suite references, and overly broad correctness claims. The document now strictly outlines active Velox backend limitations. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
