marin-ma opened a new pull request, #13109: URL: https://github.com/apache/gluten/pull/13109
Velox's Hive connector can declare required subfields on a column handle, and its Parquet map reader then materialises only the entries whose key matches, so that reading m['k'] no longer decodes every entry of every row's map. This change lets Gluten derive those declarations for map columns read by a native Parquet scan. Planner rule ScanMapKeyPruning (Velox backend, post-transform, off by default, spark.gluten.sql.columnar.backend.velox.scanMapKeyPruningEnabled): for each map-typed output of a native Parquet scan it walks up through Filter, Project and Generate nodes only. Every reference to the map on the way must be a constant-key path (GetMapValue, ElementAt on a map, GetStructField steps) or a null check, and some Project or Generate must drop the attribute, and every alias still holding map data, before the chain ends. Since an operator can only reference what its child outputs, nothing above that point can read the map and needs no analysis. Any other shape leaves the map whole: the map still live at an exchange, join, aggregate, union, write or the fragment output, a whole-map use such as size(m) or explode(m), a non-constant key, a key type without a Velox subscript form (date, decimal), or a string key whose bytes are not valid UTF-8. Null checks need no map entries and are declared only when no value path covers them; a declaration is emitted only when it lets the reader skip something. Spark's NestedColumnAliasing alias of a struct field holding the map (c.m AS _extract_m) is followed with a path prefix. Transport and native side: paths travel as a structured protobuf (RequiredSubfieldsExtension: column, then field / string key / long key elements) in the ReadRel advanced extension and are rebuilt as common::Subfield on the HiveColumnHandle, with case folded like the schema's names; declared columns that match no scan column are logged. Spark's GetMapValue is emitted as get_map_value, a Gluten-registered map-only subscript that reports canPushdown(), so a remaining filter such as m['k'].x = v extracts m["k"] instead of the whole map and does not defeat the declaration. Delta scans participate when the table uses no column mapping; BatchScan and FileSourceScan transformers carry the declaration and include it in their equality. Tests: ScanMapKeyPruningSuite runs each case with adaptive execution off and on and asserts the exact declaration of every native scan, including a direct-declaration test that the reader returns only the declared key; Delta and Hive UDF suites cover column mapping and partial generates. ## Was this patch authored or co-authored using generative AI tooling? Claude Fable 5.1 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
