clintonjrobinson opened a new issue, #26144:
URL: https://github.com/apache/datafusion/issues/26144
### Describe the bug
With `datafusion.optimizer.enable_leaf_expression_pushdown = true` (the
default), `ExtractLeafExpressions` moves a `get_field` out of an aggregate
argument into an extraction projection, and `PushDownLeafProjections` moves
that projection below a `SubqueryAlias`. The aggregate keeps an unqualified
`Column("__datafusion_extracted_N")`, so its output field is named `max(...
__datafusion_extracted_N ...)`. The optimized plan executes correctly.
`logical_plan_to_bytes` encodes it, but `logical_plan_from_bytes` fails:
decoding rebuilds the aggregate through `LogicalPlanBuilder`, which qualifies
the column as `hdr.__datafusion_extracted_N`. The aggregate's output name
changes, and the projection above it, which references the old name, no longer
resolves.
### To Reproduce
With datafusion and datafusion-proto 53.1.0 or 55.2.0 (the same failure with
a Parquet table in place of the literal):
```rust
use datafusion::prelude::SessionContext;
use datafusion_proto::bytes::{logical_plan_from_bytes,
logical_plan_to_bytes};
let ctx = SessionContext::new();
let plan = ctx.sql(
"WITH t AS (SELECT 'm1' AS id, [named_struct('name', 'Subject', 'value',
'Hello')] AS headers), \
hdr AS (SELECT id, unnest(headers) AS h FROM t) \
SELECT id, MAX(CASE WHEN lower(hdr.h['name']) = 'subject' THEN
hdr.h['value'] END) AS subject \
FROM hdr GROUP BY id",
).await?.into_optimized_plan()?;
let bytes = logical_plan_to_bytes(&plan)?;
logical_plan_from_bytes(&bytes, &ctx.task_ctx())?; // fails
```
On 53.1.0 (55.2.0 adds a list of valid fields):
```text
Schema error: No field named "max(CASE WHEN lower(__datafusion_extracted_1)
= Utf8(""subject"") THEN __datafusion_extracted_2 END)". Did you mean 'max(CASE
WHEN lower(hdr.__datafusion_extracted_1) = Utf8("subject") THEN
hdr.__datafusion_extracted_2 END)'?
```
The same SQL round-trips with `enable_leaf_expression_pushdown = false`, and
when the struct fields are projected as columns before the aggregate.
### Expected behavior
The optimized plan round-trips through datafusion-proto: either the
extracted column reference in the aggregate is qualified the way the input
schema holds it, or the aggregate expression keeps its original output name
with an alias.
### Additional context
Related to the name-resolution class in #25459 and #25448 (resolve columns
by schema index, not by name).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]