Hi everyone, I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion:
*https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting> * Anticompaction and partial-range streaming currently rewrite rows whose encoded representation already exists on disk. This consumes CPU, creates substantial heap churn and write amplification, and increases temporary disk pressure. CEP-66 proposes splitting eligible compressed SSTables by retaining contiguous runs of their existing compression chunks. Cassandra would rebuild the child SSTables' indexes and other derived components without deserializing, serializing, or recompressing their rows. "Zero-copy" here primarily means reusing the encoded bytes instead of rewriting rows. On filesystems that support range reflinks, Cassandra can also share the underlying extents meaning no new data written. Other filesystems, including ext4, would copy the already-compressed bytes and still avoid the row rewrite. The proposal is staged. It starts with an opt-in `sstablesplit --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later phases add BTI support, secondary indexes, anticompaction, and partial-range streaming. Existing implementations remain the default and provide the fallback for unsupported inputs. I'd particularly appreciate feedback on: - The retained-prefix representation and proposed Cassandra 7.0 SSTable format change - Rebuilding or conservatively deriving child metadata without decoding rows - The integrity and performance tradeoff around `Digest.crc32` generation - The staged rollout, compatibility rules, and fallback behavior - Any correctness, operational, or filesystem concerns the proposal has missed Thanks, and I look forward to the discussion. Regards, Chris Lohfink
