srielau opened a new pull request, #58413:
URL: https://github.com/apache/spark/pull/58413
<!--
Thanks for sending a pull request! Here are some tips for you:
1. If this is your first time, please read our contributor guidelines:
https://spark.apache.org/contributing.html
2. Ensure you have added or run the appropriate tests for your PR:
https://spark.apache.org/developer-tools.html
3. If the PR is unfinished, add '[WIP]' in your PR title, e.g.,
'[WIP][SPARK-XXXX] Your PR title ...'.
4. Be sure to keep the PR description updated to reflect all changes.
5. Please write your PR title to summarize what this PR proposes.
6. If possible, provide a concise example to reproduce the issue for a
faster review.
7. If you want to add a new configuration, please read the guideline first
for naming configurations in
'common/utils/src/main/scala/org/apache/spark/internal/config/ConfigEntry.scala'.
8. If you want to add or modify an error type or message, please read the
guideline first in
'common/utils/src/main/resources/error/README.md'.
-->
### What changes were proposed in this pull request?
`CharType` and `VarcharType` extend `StringType`, but equality includes the
length constraint, so `dt == StringType` and `case StringType =>` match only
unconstrained STRING. Under `spark.sql.charVarchar.standardSemantics.enabled`,
first-class CHAR/VARCHAR therefore miss several production paths that already
accept STRING.
This PR replaces those exact matches with string-family checks
(`isInstanceOf[StringType]` / `case _: StringType`) at:
- `UpCastResolution` (legacy loose Dataset upcast)
- `SpecialDatetimeValues` (foldable `CAST(... AS DATE/TIMESTAMP)`)
- `AstBuilder` LIKE ANY/ALL (`LikeAny` / `LikeAll` for foldable CHAR/VARCHAR
patterns)
- `ReusableBroadcastValueProjection` (UTF8 attributes and date-format
literals)
- `from_protobuf` / `to_protobuf` 3-arg descriptor routing
- `ProtobufSerializer` STRING / ENUM / `StringValue` fields
Exact `== StringType` is left in place where unconstrained STRING is
intentional (for example annotated-STRING `SimplifyCasts`,
`StringHelper.isPlainString`, JDBC dialect mapping, default-collation STRING).
Variant UUID-to-string is not changed: `VariantGet.checkDataType` already
rejects CHAR/VARCHAR, so that equality is unreachable and widening it would
skip length enforcement.
### Why are the changes needed?
Without the family match, CHAR/VARCHAR skip STRING-only branches: special
datetime strings are not folded, LIKE ANY/ALL miss the `LikeAny`/`LikeAll`
plan, DPP broadcast projection is skipped, protobuf descriptor paths are
dropped, and `to_protobuf` cannot serialize CHAR/VARCHAR string fields.
This is the remaining OSS exact-`StringType` slice of
[SPARK-58794](https://issues.apache.org/jira/browse/SPARK-58794).
### Does this PR introduce _any_ user-facing change?
Yes, only when `spark.sql.charVarchar.standardSemantics.enabled` is true
(default remains false). CHAR/VARCHAR then follow the same STRING branches
listed above instead of falling through. `spark.sql.legacy.doLooseUpcast` also
treats CHAR/VARCHAR like STRING for Dataset encoder upcast.
### How was this patch tested?
Added SPARK-59112 cases in:
- `EncoderResolutionSuite`
- `SpecialDatetimeValuesSuite`
- `ExpressionParserSuite`
- `DynamicPruningSubquerySuite`
- `ProtobufDescriptorFileReadSuite`
- `ProtobufSerdeSuite`
Ran those suites plus Catalyst/protobuf scalastyle locally.
### Was this patch authored or co-authored using generative AI tooling?
Generated-by: Cursor Grok 4.6
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]