Let's separate the implementation specific part from SPIP. For
implementation things, I think we can discuss details in the PR and make
sure, during the reviews, to involve people who had concerns about it.

On Fri, 18 Sept 2026 at 02:16, Gengliang Wang <[email protected]> wrote:

> Thanks for the update. I agree with separating the value domain from
> exception handling and the arithmetic implementation.
>
> I support defining DECFLOAT with signed zero, positive and negative
> infinity, and quiet NaN. These are numeric values rather than SQL NULL, and
> ppreserving them enables faithful interchange and round-tripping across
> language and engine boundaries, including between Spark and Python clients.
> A finite-only definition would lose that information and would be difficult
> to extend compatibly later.
>
> Detailed coercion rules, the execution implementation, external APIs, and
> the Parquet encoding still require further review, but they should not
> block agreement on the logical value domain.
>
> With this separation clarified, I support moving the SPIP forward.
>
> Best,
> Gengliang
>
> On Thu, Sep 17, 2026 at 7:44 AM Uroš Bojanić <[email protected]> wrote:
>
>> Hi all,
>>
>> We've had some great input and discussion in the DECFLOAT SPIP document (
>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E/edit?tab=t.0#heading=h.zarpams0nxzq),
>> so I'm writing an update here to keep the DISCUSS thread active and invite
>> further feedback and discussion regarding this SPIP.
>>
>> It's really important to separate the irreversible decisions from the
>> reversible ones. The one-way door here is the value domain - DECFLOAT
>> should carry IEEE 754's non-finite values (+Inf, -Inf, NaN,
>> signed zero), the same way DOUBLE already does. Digit count (e.g. 34 vs
>> 38) and arithmetic implementation can easily be changed later - for
>> example, digits can always be widened later and retain full backwards
>> compatibility; also, the execution kernel can change under a fixed IEEE
>> contract. However, specials are at the opposite end - we can't retrofit
>> them onto a shipped finite-only type without breaking it.
>>
>> Why this matters for big data: decimal pipelines produce unbounded and
>> undefined quantities as legitimate business states, not data bugs (e.g. an
>> uncapped credit line, a threshold that means "no maximum," a magnitude that
>> has run past any finite bound). IEEE already names & defines these
>> precisely, and DOUBLE carries them today. DECFLOAT is the base-10 DOUBLE,
>> its domain should match accordingly.
>>
>> A finite-only decimal type can't hold those values, so applications
>> disguise them (almost always as NULL). However, NULL in Spark means
>> missing/unknown, so "unbounded" and "never loaded" would become the same
>> value, and SUM, COUNT, range filters, etc. could no longer telling them
>> apart. The loss here runs one way, we can always collapse Inf/NaN to NULL
>> if we want SQL null semantics, but we can never rebuild meaning already
>> flattened into NULL. This reaches past Spark too, Parquet's decimal-float
>> work and existing C++/Rust/Python engines already emit Inf/NaN, so a
>> finite-only type can't round-trip that data faithfully across different
>> languages / engines.
>>
>> Let's agree that DECFLOAT's domain is IEEE 754's, non-finite values
>> included. Without this, DECFLOAT isn't a better DECIMAL - it's a worse
>> DOUBLE.
>>
>> Best,
>> Uroš
>>
>> On 2026/08/18 16:00:13 Uroš Bojanić wrote:
>> > Hi all,
>> >
>> > I would like to start a discussion on the SPIP to add a new Spark SQL
>> data type: DECFLOAT (IEEE 754 decimal64 / decimal128), for base-10
>> floating-point decimals with per-value exponents.
>> >
>> > JIRA ID: https://issues.apache.org/jira/browse/SPARK-58820
>> >
>> > Brief summary: DECFLOAT is a new IEEE 754 decimal floating-point type
>> (decimal64 / decimal128, i.e. DECFLOAT(16) / DECFLOAT(34)) that closes the
>> gap between DECIMAL, which caps at precision 38 with a fixed per-column
>> scale, and DOUBLE, whose binary rounding makes fractions like 0.1 inexact:
>> each value keeps its own exponent, arithmetic is decimal, and it supports
>> signed zero, Inf, and NaN. It is additive, round-trips through a proposed
>> Parquet logical type for cross-engine interop, and would ship config-gated
>> during incubation like TIME, leaving existing DECIMAL / DOUBLE behavior
>> unchanged.
>> >
>> > Additional information is available in the SPIP document:
>> >
>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E
>> >
>> > Please provide your feedback on the approach, scope, and the proposal
>> described in the document.
>> >
>> > Thank you!
>> >
>> > Best,
>> > Uroš
>> >
>> > ---------------------------------------------------------------------
>> > To unsubscribe e-mail: [email protected]
>> >
>> >
>>
>> ---------------------------------------------------------------------
>> To unsubscribe e-mail: [email protected]
>>
>>

Reply via email to