yongster commented on issue #11243: URL: https://github.com/apache/arrow-rs/issues/11243#issuecomment-5905181095
@alamb Yes — this is specifically to speed up `Float16 ↔ Float32` for AI embedding columns. Those columns are usually stored as Float16 to cut memory and I/O. A typical batch is `FixedSizeList<Float16, 768>` or `FixedSizeList<Float16, 1536>` (BERT-sized and OpenAI `text-embedding-3-small`-sized vectors). Before cosine / L2 distance or L2 normalization, the values are widened to Float32, because that is what the compute kernels take. After the math they are often narrowed back to Float16 to write the batch out. Both casts run on every batch, so this is a hot path rather than a one-off conversion. The two directions cannot fail, but they still go through the generic `unary_opt(num_cast)` path. On Apple Silicon with `fp16`, a 1024×768 batch is about 200–360 µs today and about 56–68 µs with a bulk `half` slice conversion (about 3.4×–6.3×, higher when the column has nulls). The change is only those two directions. Other numeric casts stay on the generic path. #11244 adds the embedding-sized benchmarks first. #11245 is the kernel change. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
