1fanwang commented on issue #12955: URL: https://github.com/apache/datafusion/issues/12955#issuecomment-5828820425
The `ROW_NUMBER()` rewrite seems like a good way to keep this as a composition of existing relational operators: partition each side by the set-operation columns, assign a row number within each partition, then join on both the columns and the row number. The join type and post-filter produce the `min(m, n)` and `max(m - n, 0)` multiplicities. For DataFusion, I see two possible layers for this: 1. Thread user-defined window expressions through the SQL planner, DataFrame API, and Substrait consumer, then add builder methods that construct the rewrite. 2. Add logical `Intersect`/`Except` nodes and lower them in an optimizer rule, keeping the set-operation semantics explicit until lowering. I lean toward the first option if the rewrite can be expressed cleanly with the existing `Window`, `Join`, `Filter`, and `Projection` nodes. It avoids adding another logical-plan surface, but the planner/API/Substrait plumbing needs to be weighed against the benefit of keeping the set-operation semantics visible for optimization and distributed execution. Would maintainers prefer a rewrite through existing primitives, or should this proceed as explicit logical set-operation nodes with a later lowering rule? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
