HippoBaro commented on PR #11267:
URL: https://github.com/apache/arrow-rs/pull/11267#issuecomment-5938203615

   For context, the 1%-cardinality cases have been on this branch for quite a 
while; the fixed-cardinality 20/100/400 cases were added later.
   
   > is the new string_dictionary_1pct markedly different from the other low 
cardinality tests? They all seem to have similar execution times.
   
   The main difference is that the 1%-cardinality cases include 25% nulls, 
whereas the fixed-cardinality cases contain none. My intention was consistency 
with the existing benchmarks: general-purpose cases such as primitive, bool, 
string, and list_primitive use 25% null density, with `_non_null` variants 
explicitly covering non-null data.
   
   The 1%-cardinality cases also form a consistent set of low-cardinality 
counterparts to the high-cardinality dictionary benchmarks across `Int32`, 
`Int64`, `Float64`, strings, and `Decimal128`, rather than being 
string-specific cases.
   
   Perhaps we should append `_non_null` to the fixed-cardinality string 
benchmark names to make that distinction clearer and align them with that 
naming convention? 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to