yongster opened a new issue, #11243:
URL: https://github.com/apache/arrow-rs/issues/11243
### Is your feature request related to a problem or challenge?
`Float16 -> Float32` and `Float32 -> Float16` go through the generic numeric
cast. The safe path calls `unary_opt(num_cast)`, so every element builds an
`Option`, and a null bitmap is rebuilt even though these conversions cannot
fail.
That is the cast used for embedding columns such as `FixedSizeList<Float16,
768>` and `FixedSizeList<Float16, 1536>` when the values are widened to
`Float32` for distance or normalization, or narrowed back to `Float16` for
storage.
`arrow-cast` already depends on `half`. `half::slice::HalfFloatSliceExt`
converts a whole slice with `convert_to_f32_slice` / `convert_from_f32_slice`.
On targets that enable `fp16` or `f16c`, that uses the hardware conversion. The
generic path does not.
This is not #10955. That issue is about infallible integer and float
widening via `AsPrimitive`. It excludes `Float16` because `as` casts do not
cover `half::f16`.
### Describe the solution you'd like
Special-case the two directions:
- Fill a new values buffer with `HalfFloatSliceExt`.
- Clone the input null buffer instead of rebuilding validity.
- Leave every other numeric cast on the generic path.
- Keep the first version on safe `Vec` initialization. Do not introduce
`unsafe` to skip zeroing.
Both directions are infallible (`f16` widens exactly, `f32` rounds to an
`f16`), so `CastOptions { safe: false }` should keep succeeding and match the
safe result.
Null slots are not logical values. Converting the whole values buffer,
including null slots, and then reusing the original validity is enough. Tests
should compare validity and valid elements, not the physical bytes stored in
null slots.
I remeasured on Apple Silicon (`target_feature=fp16`), rustc 1.97, release +
thin LTO, median of 51 samples of 20 calls, against `fa337f843`. The prototype
is the slice conversion above, wrapped in the same `FixedSizeList` cast the
embedding path uses.
| Shape | Direction | Current | Slice prototype | Speedup |
|---|---|---:|---:|---:|
| 1024 × 768 | f16 → f32 | 218 µs | 64 µs | 3.40× |
| 1024 × 768 | f32 → f16 | 206 µs | 56 µs | 3.67× |
| 1024 × 1536 | f16 → f32 | 421 µs | 123 µs | 3.42× |
| 1024 × 1536 | f32 → f16 | 410 µs | 111 µs | 3.69× |
| 1024 × 768, 10% null | f16 → f32 | 366 µs | 68 µs | 5.37× |
| 1024 × 768, 10% null | f32 → f16 | 354 µs | 56 µs | 6.30× |
Correctness against the current `cast`: all 65,536 `f16` bit patterns, and
1,000,000 deterministic `f32` bit patterns, matched bit for bit. Fixed-size
list output compared equal, including the nullable child case.
### Describe alternatives you've considered
- Routing these pairs through `PrimitiveArray::unary` instead of
`unary_opt`. That removes the `Option` and the validity rebuild, but still
converts one value at a time and does not use the `half` slice kernel.
- Enabling the `std` feature of `half` so CPUs without the compile-time
`fp16` / `f16c` target feature can detect it at runtime. Not required for this
change: `half` already selects the hardware path when the target has the
feature, which is how the numbers above were measured.
- An uninitialized output buffer. That can avoid the zero fill, but it needs
`unsafe`. Worth a later pass only if the fill shows up in a profile.
### Additional context
`arrow/benches/cast_kernels.rs` has no `Float16` case today. Per the
contributing guide, the new embedding-sized benchmarks should land in a
separate PR first so the benchmark runner has a baseline before the kernel
change.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]