Thanks, Branimir. The shared-index approach is an attractive option
especially for transient local operations such as anticompaction.

The tradeoff appears to be where the complexity lives. Reusing the original
primary index makes slice creation cheaper, but its positions remain in the
parent Data.db coordinate space. Readers and tools must therefore
understand that the Data.db component is a slice and translate those
positions before accessing the local file.

The current proposal instead pays that cost once during splitting. It
rebases Index.db positions, slices CompressionInfo.db, and rebuilds or
apportions the child’s derived metadata. This preserves the usual idea that
each SSTable is a self-contained artifact whose components share one
coordinate space and lifecycle. Index reconstruction is so cheap its
comparatively free relative to reading or copying Data.db, while also
producing child-specific summaries, Bloom filters, and statistics (100's of
ms range on *huge* sstables)

For durable outputs, I prefer keeping that complexity at creation time.
Sharing one physical index file among several logical SSTables would
introduce additional lifecycle and bookkeeping cases around deletion,
snapshots, backup and restore, import, and third-party tooling. Independent
files that happen to share filesystem extents are less concerning because
they retain normal component ownership semantics.

That said, the DataStax approach is worth benchmarking, and it may be a
better fit for BTI or another format where slicing is designed in as a
first-class property.

> Separately, being able to easily slice and dice files without looking
inside them is a key consideration in the file format we are working on for
CEP-57.

Agreed,nCEP-57 seems like the right place to make sliceability a native
format property. My goal here is narrower: provide the capability for BIG
SSTables and current formats in the meantime without permanently
introducing slice-coordinate awareness throughout BIG’s read path. BIG will
remain in use for some time, and this mechanism can deliver most of the
benefit with relatively contained changes.

On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> wrote:

> Hello Chris,
>
> A while back we implemented a similar approach for DSE's version of
> zero-copy streaming. The main difference between our approach and yours is
> that we decided not to split the primary index files and instead use them
> as they are, with filtering based on the start and end key of the section.
> In the context of local operations like anticompaction, the latter may be a
> better approach as one can share the index files between all resulting
> sections (and, of course, no index reconstruction is necessary).
>
> This code is not currently part of the Apache codebase, but DataStax's
> open source fork includes support for reading these files, whose
> implementation (commit
> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422)
> may have some ideas you can use in designing your solution.
>
> Separately, being able to easily slice and dice files without looking
> inside them is a key consideration in the file format we are working on for
> CEP-57.
>
> Regards,
> Branimir
>
> On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
> [email protected]> wrote:
>
>> Really nice idea Chris!
>>
>> One thing I think it’s worth flagging for the proposal:
>>
>> this CEP introduces a new SSTable major version. That's a relevant change
>> for anything outside nodetool/the core read path that parses SSTables
>> directly. e.g. analytics library uses the concept of per major version
>> bridge to deserialize.
>>
>> With this approach, current bridges don't have a notion of "data doesn't
>> start at logical offset zero” - they assume a chunk's first partition
>> begins the SSTable's data - which is fine (this is a new version!). It'd
>> help if the CEP explicitly calls out that third-party/off-node SSTable
>> readers are a compatibility surface here, not just in-process Cassandra
>> binaries.
>>
>> Bernardo
>>
>> *From: *Chris Lohfink <[email protected]>
>> *Date: *Monday, 14 September 2026 at 21:23
>> *To: *[email protected] <[email protected]>
>> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
>>
>> Hi everyone,
>>
>> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion:
>>
>>
>> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
>> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting>
>>
>>
>> Anticompaction and partial-range streaming currently rewrite rows whose
>> encoded representation already exists on disk. This consumes CPU, creates
>> substantial heap churn and write amplification, and increases temporary
>> disk pressure.
>>
>> CEP-66 proposes splitting eligible compressed SSTables by retaining
>> contiguous runs of their existing compression chunks. Cassandra would
>> rebuild the child SSTables' indexes and other derived components without
>> deserializing, serializing, or recompressing their rows.
>>
>> "Zero-copy" here primarily means reusing the encoded bytes instead of
>> rewriting rows. On filesystems that support range reflinks, Cassandra can
>> also share the underlying extents meaning no new data written. Other
>> filesystems, including ext4, would copy the already-compressed bytes and
>> still avoid the row rewrite.
>>
>> The proposal is staged. It starts with an opt-in `sstablesplit
>> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later
>> phases add BTI support, secondary indexes, anticompaction, and
>> partial-range streaming. Existing implementations remain the default and
>> provide the fallback for unsupported inputs.
>>
>> I'd particularly appreciate feedback on:
>>
>> - The retained-prefix representation and proposed Cassandra 7.0 SSTable
>> format change
>> - Rebuilding or conservatively deriving child metadata without decoding
>> rows
>> - The integrity and performance tradeoff around `Digest.crc32` generation
>> - The staged rollout, compatibility rules, and fallback behavior
>> - Any correctness, operational, or filesystem concerns the proposal has
>> missed
>>
>> Thanks, and I look forward to the discussion.
>>
>> Regards,
>> Chris Lohfink
>>
>

Reply via email to