cdelmonte-zg commented on issue #16919:
URL: https://github.com/apache/datafusion/issues/16919#issuecomment-5886638343

   Hi @zheniasigayev , I looked into this on `e1aa7d956` and noticed a problem 
with the generator in the gist: it concatenates entities in encounter order and 
sorts some of them by `col_2` descending. So the files do not satisfy `WITH 
ORDER (col_1 ASC, col_2 ASC)`.
   
   I ran the original generator and found this inversion in the first file, at 
row index 42030:
   
   ```text
   previous: ('5564918-58-C--Z16B9723620', 1790595853274)
   current:  ('1483923-38-A--Z13C1375906', 1790590344849)
   ```
   
   Are the production files sorted by both columns? I changed the generator to 
sort each complete file before writing it and continued testing with those 
files.
   
   There is still an issue worth exploring with sorted input. With 
`split_file_groups_by_statistics = true`, overlapping files can require more 
groups than `target_partitions`. `ListingTable` rejects those extra groups, and 
the scan loses its advertised ordering.
   
   I temporarily removed that check. With 15 sorted files, 600,000 rows total 
and target 10, the plan used 15 groups, switched both aggregates to 
`PartiallySorted([0, 1])`, and no longer needed `SortExec`. The original plan 
ran out of memory with a `fair` pool of both 128 MB and 256 MB; the modified 
plan completed at both limits.
   
   But simply removing the check isn't enough. With 1,000 files and target 2, 
it hit the open-file limit. After raising that limit, it completed but was 
about 9% slower and used almost three times the peak process RSS.
   
   Would the generator with each file fully sorted match your real data? And 
would allowing a limited number of extra file groups be worth pursuing?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to