Thanks for updating the proposal and working through the feedback. I
support adding decimal floating-point support to Spark.

The revised proposal now specifies an in-tree, pure-Java implementation
with no native runtime dependency. This appears to address the
architectural concern behind Dongjoon’s original -1. Dongjoon, could you
confirm whether this addresses your concern about having a safe Java
execution path?

Thanks,

On Mon, Sep 21, 2026 at 3:40 PM Hyukjin Kwon <[email protected]> wrote:

> Let's separate the implementation specific part from SPIP. For
> implementation things, I think we can discuss details in the PR and make
> sure, during the reviews, to involve people who had concerns about it.
>
> On Fri, 18 Sept 2026 at 02:16, Gengliang Wang <[email protected]> wrote:
>
>> Thanks for the update. I agree with separating the value domain from
>> exception handling and the arithmetic implementation.
>>
>> I support defining DECFLOAT with signed zero, positive and negative
>> infinity, and quiet NaN. These are numeric values rather than SQL NULL, and
>> ppreserving them enables faithful interchange and round-tripping across
>> language and engine boundaries, including between Spark and Python clients.
>> A finite-only definition would lose that information and would be difficult
>> to extend compatibly later.
>>
>> Detailed coercion rules, the execution implementation, external APIs, and
>> the Parquet encoding still require further review, but they should not
>> block agreement on the logical value domain.
>>
>> With this separation clarified, I support moving the SPIP forward.
>>
>> Best,
>> Gengliang
>>
>> On Thu, Sep 17, 2026 at 7:44 AM Uroš Bojanić <[email protected]> wrote:
>>
>>> Hi all,
>>>
>>> We've had some great input and discussion in the DECFLOAT SPIP document (
>>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E/edit?tab=t.0#heading=h.zarpams0nxzq),
>>> so I'm writing an update here to keep the DISCUSS thread active and invite
>>> further feedback and discussion regarding this SPIP.
>>>
>>> It's really important to separate the irreversible decisions from the
>>> reversible ones. The one-way door here is the value domain - DECFLOAT
>>> should carry IEEE 754's non-finite values (+Inf, -Inf, NaN,
>>> signed zero), the same way DOUBLE already does. Digit count (e.g. 34 vs
>>> 38) and arithmetic implementation can easily be changed later - for
>>> example, digits can always be widened later and retain full backwards
>>> compatibility; also, the execution kernel can change under a fixed IEEE
>>> contract. However, specials are at the opposite end - we can't retrofit
>>> them onto a shipped finite-only type without breaking it.
>>>
>>> Why this matters for big data: decimal pipelines produce unbounded and
>>> undefined quantities as legitimate business states, not data bugs (e.g. an
>>> uncapped credit line, a threshold that means "no maximum," a magnitude that
>>> has run past any finite bound). IEEE already names & defines these
>>> precisely, and DOUBLE carries them today. DECFLOAT is the base-10 DOUBLE,
>>> its domain should match accordingly.
>>>
>>> A finite-only decimal type can't hold those values, so applications
>>> disguise them (almost always as NULL). However, NULL in Spark means
>>> missing/unknown, so "unbounded" and "never loaded" would become the same
>>> value, and SUM, COUNT, range filters, etc. could no longer telling them
>>> apart. The loss here runs one way, we can always collapse Inf/NaN to NULL
>>> if we want SQL null semantics, but we can never rebuild meaning already
>>> flattened into NULL. This reaches past Spark too, Parquet's decimal-float
>>> work and existing C++/Rust/Python engines already emit Inf/NaN, so a
>>> finite-only type can't round-trip that data faithfully across different
>>> languages / engines.
>>>
>>> Let's agree that DECFLOAT's domain is IEEE 754's, non-finite values
>>> included. Without this, DECFLOAT isn't a better DECIMAL - it's a worse
>>> DOUBLE.
>>>
>>> Best,
>>> Uroš
>>>
>>> On 2026/08/18 16:00:13 Uroš Bojanić wrote:
>>> > Hi all,
>>> >
>>> > I would like to start a discussion on the SPIP to add a new Spark SQL
>>> data type: DECFLOAT (IEEE 754 decimal64 / decimal128), for base-10
>>> floating-point decimals with per-value exponents.
>>> >
>>> > JIRA ID: https://issues.apache.org/jira/browse/SPARK-58820
>>> >
>>> > Brief summary: DECFLOAT is a new IEEE 754 decimal floating-point type
>>> (decimal64 / decimal128, i.e. DECFLOAT(16) / DECFLOAT(34)) that closes the
>>> gap between DECIMAL, which caps at precision 38 with a fixed per-column
>>> scale, and DOUBLE, whose binary rounding makes fractions like 0.1 inexact:
>>> each value keeps its own exponent, arithmetic is decimal, and it supports
>>> signed zero, Inf, and NaN. It is additive, round-trips through a proposed
>>> Parquet logical type for cross-engine interop, and would ship config-gated
>>> during incubation like TIME, leaving existing DECIMAL / DOUBLE behavior
>>> unchanged.
>>> >
>>> > Additional information is available in the SPIP document:
>>> >
>>> https://docs.google.com/document/d/1qnLXm0ldHSwPSJs_Q5SzGyyTKMMEqeX5rm8Fh4rvV3E
>>> >
>>> > Please provide your feedback on the approach, scope, and the proposal
>>> described in the document.
>>> >
>>> > Thank you!
>>> >
>>> > Best,
>>> > Uroš
>>> >
>>> > ---------------------------------------------------------------------
>>> > To unsubscribe e-mail: [email protected]
>>> >
>>> >
>>>
>>> ---------------------------------------------------------------------
>>> To unsubscribe e-mail: [email protected]
>>>
>>>

Reply via email to