kosiew commented on code in PR #25461: URL: https://github.com/apache/datafusion/pull/25461#discussion_r4204575847
########## datafusion/sqllogictest/test_files/aggregate.slt: ########## @@ -404,6 +404,25 @@ GROUP BY g 1 1 2 NULL +# A DISTINCT aggregate may take the distinct argument more than once. +# Group 1 aggregates over the distinct values 1.0 and 2.0, group 2 over the +# single distinct value 3.0, which leaves the sample covariance undefined. +query IR rowsort +SELECT g, corr(DISTINCT x, x) +FROM (VALUES (1, 1.0), (1, 1.0), (1, 2.0), (2, 3.0), (2, 3.0)) AS t(g, x) +GROUP BY g +---- +1 1 +2 NULL + +query IR rowsort +SELECT g, covar_samp(DISTINCT x, x) +FROM (VALUES (1, 1.0), (1, 1.0), (1, 2.0), (2, 3.0), (2, 3.0)) AS t(g, x) Review Comment: Optional: could we also add repeated NULL rows and an all-NULL group here? That would cover how the repeated-argument DISTINCT rewrite interacts with NULL grouping and the aggregate's NULL exclusion, although the current tests already cover the reported bug. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
