Hi everyone,

Adding details about the DBP kernel vs. PFOR-delta comparison below.

Antoine recently asked if we could use Parquet’s DELTA_BINARY_PACKED (DBP)
kernel directly for the delta payload. Alkis subsequently asked whether
DBP’s existing decoder could be optimized, so I measured both.

Here are the three configurations tested: (all numbers on AWS graviton 4
machines)

   1.

   *PFOR-PatchedDelta*: PFOR’s current delta layout, using one frame and
   bit width per 1,024-value vector, with outliers stored as exceptions.
   2.

   *PFOR-DBPDelta*: A canonical DBP stream restarted inside each
   1,024-value PFOR vector.
   3.

   *PFOR-DBPDelta optimized*: The same DBP bytes, but with equal-width
   miniblock coalescing and a SIMD prefix-sum decoder.

*Delta-beneficial datasets (18 cases):*

                         Compression ratio    Decode GB/s
PFOR-PatchedDelta               10.072           9.852
PFOR-DBPDelta                    6.536           3.763
PFOR-DBPDelta optimized          6.536           5.964

*Datasets where delta mode is not beneficial (36 cases):*

                         Compression ratio    Decode GB/s
PFOR-PatchedDelta                2.829           7.847
PFOR-DBPDelta                    2.771           2.657
PFOR-DBPDelta optimized          2.771           4.861

*Analysis:* The DBP optimizations substantially narrow the decompression
gap. Nevertheless, PFOR-PatchedDelta remains:

   -

   *35.1% smaller* and *1.65x faster* to decode on delta-beneficial
   datasets.
   -

   *2.1% smaller* and *1.61x faster* on the non-beneficial datasets.

The remaining difference appears to be structural. PatchedDelta uses one
bit width and generally one unpack invocation for an entire 1,024-value
vector, followed by exception patching and one prefix sum. DBP divides the
same vector into blocks and miniblocks, each potentially having a different
width. The optimized decoder removes much of the avoidable overhead, but it
must still parse these boundaries and execute a data-dependent number of
unpack runs.

*Based on these results:*

   -

   *Retain PFOR-PatchedDelta:* I recommend retaining PFOR-PatchedDelta as
   the current PFOR delta payload rather than replacing it with DBP. We can
   continue evaluating FastLanes and interleaved layouts separately through
   the layout discriminator.
   -

   *Separate FastLanes:* I propose that we continue the current PFOR work
   using the sequential bit-packing layout and keep the FastLanes work
   separate. The PFOR format can retain a layout discriminator so that we have
   room to introduce additional layouts later—for example, the FastLanes
   layout for delta values, and potentially an interleaved layout for the
   regular bit-packing kernels—without making those layouts part of the
   initial PFOR proposal.

What do people think on the above recommendation?

I will incorporate these results into the main PFOR design document this
week. Happy to answer questions or run additional comparisons.

Best, Prateek


On Tue, Sep 22, 2026 at 11:28 AM PRATEEK GAUR <[email protected]> wrote:

> Apologies for the brevity, typing on phone.
>
> @alkis : embedding DBP (delta binary packed ) into PFOR is feasible. Both
> build on same principles. The reason I wanted it separate was to have
> flexibility in deciding a new layout without being slowed down by DBP spec.
> I can think more and get back on this.
>
> @Andrew Lamb <[email protected]>  : DBP -- delta binary packed
> (encoding)
>
>
> Best
> Prateek
>
>
>
> On Tue, Sep 22, 2026, 10:43 AM Andrew Lamb <[email protected]> wrote:
>
>> Sorry for my ignorance -- what does DBP stand for?
>>
>> On Tue, Sep 22, 2026 at 1:24 PM Alkis Evlogimenos via dev <
>> [email protected]> wrote:
>>
>> > I left a comment in the doc. Posting here as well:
>> >
>> > 1. Can we compare PFOR-delta vs DBP?
>> > 2. Is there a design where we can embed DBP blocks between PFOR blocks
>> and
>> > get the same benefits without opening the writer/encoder to more
>> choices?
>> >
>> >
>> > On Tue, Sep 8, 2026 at 5:07 PM PRATEEK GAUR <[email protected]> wrote:
>> >
>> > > Hi all,
>> > >
>> > > Since the earlier PFOR thread I have kept measuring and improved
>> decoding
>> > > speed
>> > > significantly, added a delta mode, and evaluated the encoding on real
>> > > dictionary index runs.
>> > > The detailed write-up for all of this has been added to the original
>> doc
>> > > itself : [1] .
>> > >
>> > > First, decode.
>> > > With code changes and with no encoded byte changing, made PFOR
>> > > about 40% faster to decode at -O3.
>> > >
>> > > Second, added delta mode.
>> > > Certain data distributions definitely benefit from delta encoding and
>> > > wanted to explore this
>> > > capability on top of PFOR's blocking + patching scheme and results
>> > looking
>> > > quite promising
>> > > for the data distributions that can take advantage of it. Here I
>> added it
>> > > as a per block decision
>> > > mode which selects delta mode if the values are close together. I see
>> a
>> > > comparable compression
>> > > ratio to delta bit pack hybrid (same idea). And pretty good
>> decompression
>> > > speed. Added details to
>> > > the same document. Code for the same is in [2].
>> > >
>> > > Third, PFOR for dictionary index runs.
>> > > Based on the characteristics of PFOR and the dictionary index runs I
>> was
>> > > hopeful
>> > > that I'll get good results with PFOR on dictionary indices. So I did
>> an
>> > > evaluation of PFOR for dictionary
>> > > index runs and added the results on the same to the document. I do see
>> > > value in using PFOR for
>> > > dictionary index encoding.
>> > >
>> > > Looking for feedback from the team.
>> > >
>> > > Best
>> > > Prateek
>> > >
>> > > [1]
>> > >
>> > >
>> >
>> https://docs.google.com/document/d/1ZZOtxmq6K8pNU0npijfSglTJVkspXL5GLDKPSGj9HlA/
>> > > [2] https://github.com/apache/arrow/pull/50088 c++ (pfor+frame only)
>> > > [3] https://github.com/apache/arrow/pull/51150 c++ (pfor + dynamic
>> frame
>> > > or
>> > > delta)
>> > > [4] https://github.com/apache/arrow-rs/pull/10977 rust
>> > >
>> > > On Mon, Jul 13, 2026 at 8:31 AM PRATEEK GAUR <[email protected]>
>> wrote:
>> > >
>> > > > Hi team,
>> > > >
>> > > > Just wanted to resurface the doc in your email folders, looking
>> forward
>> > > to
>> > > > some suggestions.
>> > > > I'll work towards adding a few more datasets to the comparison in
>> the
>> > > > coming days.
>> > > >
>> > > > Best
>> > > > Prateek
>> > > >
>> > > > On Wed, Jul 1, 2026 at 9:01 AM PRATEEK GAUR <[email protected]>
>> > wrote:
>> > > >
>> > > >> Hi Antoine,
>> > > >>
>> > > >> Apologies for the delay. Got some time to work on it and updated
>> the
>> > doc
>> > > >> with the requested numbers.
>> > > >>
>> > > >>    - Obtained the numbers for ZSTD again. I agree they were
>> incorrect.
>> > > >>    - Fixed the algorithm to use decompressed size/data
>> > > >>    - Added BSS + ZSTD to the evaluation metrics
>> > > >>    - Added BSS + Lz4 to the evaluation metrics
>> > > >>
>> > > >>
>> > > >> Updated doc :
>> > > >>
>> > >
>> >
>> https://docs.google.com/document/d/1ZZOtxmq6K8pNU0npijfSglTJVkspXL5GLDKPSGj9HlA/edit?tab=t.0#heading=h.uzgoevv9ajp
>> > > >> Will be pushing the branch/code for the same today.
>> > > >>
>> > > >> Best
>> > > >> Prateek
>> > > >>
>> > > >> On Wed, Dec 10, 2025 at 11:30 AM Antoine Pitrou <
>> [email protected]>
>> > > >> wrote:
>> > > >>
>> > > >>>
>> > > >>> Hi,
>> > > >>>
>> > > >>> I looked at the doc and the stated decompression speeds for ZSTD
>> look
>> > > >>> highly irrealistic.
>> > > >>>
>> > > >>> I think what happens is that you are computing decompression speed
>> > as:
>> > > >>>
>> > > >>>    size of compressed data / time to decompress
>> > > >>>
>> > > >>> while you should really compute it as:
>> > > >>>
>> > > >>>    size of uncompressed data / time to decompress
>> > > >>>
>> > > >>> Otherwise you're simply making ZSTD look miserable because it
>> > > compresses
>> > > >>> so well.
>> > > >>>
>> > > >>> Also, I think you should also add BYTE_STREAM_SPLIT + ZSTD into
>> the
>> > mix
>> > > >>> (and possible BYTE_STREAM_SPLIT + LZ4 if you're going to evaluate
>> > LZ4).
>> > > >>>
>> > > >>> Regards
>> > > >>>
>> > > >>> Antoine.
>> > > >>>
>> > > >>>
>> > > >>> Le 06/12/2025 à 23:24, PRATEEK GAUR a écrit :
>> > > >>> > Hi team,
>> > > >>> >
>> > > >>> > We wanted to share performance numbers for one of the candidate
>> > > >>> encodings
>> > > >>> > we have been discussing for parquet, for numeric compression.
>> > > >>> >
>> > > >>> > Following doc, PFOR : Encoding
>> > > >>> > <
>> > > >>>
>> > >
>> >
>> https://docs.google.com/document/d/1ZZOtxmq6K8pNU0npijfSglTJVkspXL5GLDKPSGj9HlA/edit?tab=t.0
>> > > >>> >,
>> > > >>> > talks through numbers on compression speed, compression ratio
>> and
>> > > >>> > decompression speed for columns values from the clickbench data.
>> > > Based
>> > > >>> on
>> > > >>> > the numbers PFOR gives superior decompression speed compared to
>> > > >>> > DELTABITPACK and RLEBITPACKHYBRID and fares better on average in
>> > > >>> > compression ratio. Comparison has also been made with ZSTD
>> > > compression
>> > > >>> in
>> > > >>> > the above doc.
>> > > >>> >
>> > > >>> > We plan to expand to a few more datasets and work on a
>> prototype in
>> > > >>> arrow
>> > > >>> > cpp code in early Jan.
>> > > >>> > Looking forward to the feedback from the group.
>> > > >>> >
>> > > >>> > Best
>> > > >>> > Prateek
>> > > >>> >
>> > > >>>
>> > > >>>
>> > > >>>
>> > >
>> >
>>
>

Reply via email to