Dongjoon, You make many valid points, and we’ll address them. Some I do not understand however, and as the author of the Decimal128 PR let me ask for clarification: It is my understanding that it is frowned upon to roll in “unused APIs” (i.e. dead code). The PR remains in draft mode because this SPIP has not passed yet. The PR’s initial purpose is to support THIS SPIP. By the same token, there ought not be a DecFloatType without this SPIP. Are you advocating: a) Roll in this library as a precondition to approve the SPIP b) Make it clear that this is the specific library being used, c) Sssert that there will be “some” java native library. Whether that is BigDecimal based, libbid, or something else - possibly home grown.
I’m also confused by the asserted dependency on Parquet and the transport layer. Aside from the obvious chicken-egg problem I wonder why the in-memory layout used by the DECFLOAT type in Spark must be aligned with the on-disk layout of Parquet. TIMESTAMP nano being one example where that is not case. Certainly it is “nice” to just copy through, but, in my opinion semantics are way more important. Which is why we have been focusing on this. Cheers Serge On Sep 22, 2026, at 8:29 AM, Thomas Kissinger via dev <[email protected]> wrote: Hi all, Thanks, Uroš, for revising the proposal and for all the work invested. I support adding a decimal-floating-point type to Spark, and I agree with Dongjoon that the proposal is not yet ready for another vote. My concern is that IEEE BID seems to be a predetermined choice rather than the result of comparing alternatives against Spark's requirements. IEEE BID is standardized, but its implementation ecosystem is very limited: neither Java nor Python provides a native BID value type. The proposed libbid-based arithmetic stops at 34 digits, leaving the established 38-digit SQL precision unsupported. Wider formats would require significant additional implementation work, not just selecting a larger precision. Porting libbid to Java reuses established algorithms and test vectors, but does not automatically transfer the original implementation's maturity. Roughly 20K lines of newly ported arithmetic still create a substantial obligation for correctness, review, and long-term maintenance. Java and Python already have mature arbitrary-precision decimal arithmetic. I would first evaluate a Spark-owned wrapper around BigDecimal for finite arithmetic, with explicit handling of special values and error semantics. This could preserve wider precision and much less custom arithmetic code—a particularly important tradeoff for the main use case of decimal-floating point types: financial workloads. Best, Thomas On Tue, Sep 22, 2026 at 4:17 PM Dongjoon Hyun <[email protected]<mailto:[email protected]>> wrote: Thank you for the revised proposal and for the follow-up. Let me start with my overall assessment, since I said the same thing on the vote thread in August: this SPIP is still not mature enough to go to a vote. The revision fixes one architectural point, but the document still carries [OPEN] markers on its public API surface, assumes a Parquet encoding that the Parquet community has explicitly deferred, and leans on a config gate and a timeline whose precedent does not support them. Details below. On DB's specific question: yes, Revision 1 addresses the concern behind my -1. The arithmetic path no longer depends on a native library, and the fallback story is now a pure-Java, in-tree implementation (Appendix E, item 6; SPARK-59111). That was the blocking factor for me, and it is resolved at the SPIP level. One follow-up on the same item: it still says native kernels "remain an optional optimization under the same IEEE contract". I'd like the SPIP to state plainly that the pure-Java path is the only arithmetic path in v1, and that any native or accelerator path is out of scope and would need its own SPIP. Otherwise the Java implementation risks being demoted to a reference path that only the test suite exercises. That said, I want to be precise about what exists today. The only code on the table is PR #58410, which adds a ported BID arithmetic library under a non-Spark package (org.bidfp): about 21K lines of Java, 10K lines of tests, and 129K lines of Intel test vectors. By its own description it "does not yet connect it to a user-facing Spark API or execution path". There is still no DecFloatType, no parser or Catalyst integration, no cast or coercion code, and no data source path. In other words, the SPIP has been revised on paper, but the implementation that would let us judge the pure-Java claim in practice has not been provided. The PR itself has not been reviewed yet either. For that reason I don't think the library should be merged as a standalone module with no consumer in the tree. It should land together with the first code that actually uses it, so that it is reviewed and exercised as part of the feature rather than parked as dead code. The SPIP-level items I'd want settled before a new vote are contracts rather than implementation details: 1. Parquet. Appendix D assumes a BID little-endian FIXED_LEN_BYTE_ARRAY of 8/16 bytes. That is not where the parquet-dev thread is. As of the Sep 5 and Sep 21 messages, the points both sides agree on are: precision parameterized rather than fixed at 38, with narrower encodings allowed and a path to wider ones; coverage of the full finite decimal128 range; at minimum +/-Inf and a canonical NaN; and, explicitly, that the physical representation is to be selected only after public benchmarks and validation vectors. Canonicalization and sNaN/-0 handling are still open, and a joint proposal has not been written. I'd suggest the SPIP make Parquet persistence conditional on the Parquet logical type being adopted, and state that Spark will not ship a private encoding in the meantime. Russell raised the alignment question on the vote thread and I share it. 2. External types and Arrow. Appendix B (Java/Scala external type, JDBC) and the Arrow transport question are still marked [OPEN]. These define what Row.get, Spark Connect, and Arrow-based collect in PySpark return, so they are public API. Ian's -0 asked for the Arrow half of this, and the revision does not answer it yet. 3. Gate and timeline. The SPIP cites the TIME type as the incubation pattern and bases the 9-12 month estimate on it. TIME is the worst possible example to lean on. Its SPIP passed on 2025-02-26. The type was developed without a gate for about nine months, and spark.sql.timeType.enabled was added on 2025-12-10, six days before the 4.1.0 release and during the RC phase, because support was incomplete (SPARK-54609). Nineteen months after the vote it is still off by default everywhere: in 4.1.0, in 4.2.0, in branch-4.3 whose RC1 was cut on 2026-08-31, in branch-4.x (4.4.0-SNAPSHOT), and on master, which is already Spark 5.0.0-SNAPSHOT. Coverage commits for TIME are still landing this month. So a config gate by itself is not a mitigation; it defers the risk across an entire major version line. DECFLOAT must not follow that precedent. We should not accept a second data type that lands piecemeal behind a flag with no defined end state, and I would treat repeating the TIME pattern as a reason to block rather than as an established practice to cite. Concretely, I'd ask the SPIP to define measurable exit criteria for the gate (function list, data sources, clients, plus the Parquet and Arrow contracts above), name who owns reaching them, and revisit the estimate, since DECFLOAT is strictly larger than TIME. A few smaller design points in the current text should also be fixed: - Coercion note (3) claims that lossy DECIMAL(35..38) -> DECFLOAT(34) widening is not implicit in ANSI mode, "consistent with ANSI coercion". Spark's ANSI mode already widens DECIMAL(38) + DOUBLE implicitly to DOUBLE, so the note describes a rule Spark does not have. The precedence also stops being a total order, so the result of DOUBLE + DECIMAL(38) + DECFLOAT(34) would depend on operand order. - Appendix A declares DecFloatType as an AtomicType. Arithmetic and aggregates such as SUM and ABS accept NumericType, so it needs to be a FractionalType like DecimalType and DoubleType. - Ordering is described as IEEE totalOrder, which distinguishes -0 from +0. Spark's DOUBLE treats -0.0 = 0.0 and NaN = NaN for equality, grouping, and joins, and sorts NaN after +Inf. DECFLOAT should follow those conventions, or equality and ORDER BY will disagree. Thanks, Dongjoon. --------------------------------------------------------------------- To unsubscribe e-mail: [email protected]<mailto:[email protected]> -- THOMAS KISSINGER Staff Software Engineer MOBILE +49 174-2195270 EMAIL [email protected]<mailto:[email protected]> [https://www.snowflake.com/wp-content/themes/snowflake/img/[email protected]] Snowflake Inc. 135 Constitution Drive Menlo Park, CA 94025, USA
