Thank you all.

The vibe from this thread I am getting is "I can proceed".

Regards

On Tue, Sep 8, 2026 at 3:20 PM Isaac Reath <[email protected]> wrote:
>
> Making it overridable via JMX and a persistent via config, similar to how we 
> did for compaction in  https://issues.apache.org/jira/browse/CASSANDRA-21169, 
> would go a long way towards making it safer to roll out per node as well.
>
>
>
> On Tue, Sep 8, 2026 at 6:33 AM Štefan Miklošovič <[email protected]> 
> wrote:
>>
>> per node
>>
>> nodetool sjk mx -b org.apache.cassandra.db:type=StorageService -ms -f
>> FlushCompression -v table
>>
>> per table
>>
>> nodetool sjk mx -b
>> "org.apache.cassandra.db:type=Tables,keyspace=ks,table=tbl" -ms -f
>> FlushCompression -v none
>>
>> On Tue, Sep 8, 2026 at 3:07 AM Jaydeep Chovatia
>> <[email protected]> wrote:
>> >
>> > >I don't understand why it is in yaml only, we have to restart a node.
>> > yaml option allows the capability to incrementally roll out the feature. 
>> > For example, when we change the compression, we don't want to apply it to 
>> > all node clusters at the same time. Can you please explain how we can 
>> > incrementally roll out the feature as a table param?
>> >
>> >
>> > Jaydeep
>> >
>> > On Mon, Sep 7, 2026 at 8:57 AM Štefan Miklošovič <[email protected]> 
>> > wrote:
>> >>
>> >> Yes, you already said it. Also some write patterns, like TWCS* with
>> >> long compaction windows, the compaction might never get a chance to
>> >> perform early enough so data will stay on disk compressed with
>> >> something they would not be compressed with if it did. Also we can't
>> >> satisfy both holding a high-ingest table that must not stall flushes
>> >> and a cold table that wants its compression ratio immediately.
>> >>
>> >> We have also introduced CEP-49 where we enabled the integration of
>> >> hardware-backed acceleration, which might compress faster to such an
>> >> extent that setting flush compression to it would be something they
>> >> could not have entertained before for some of their tables.
>> >>
>> >> I don't understand why it is in yaml only, we have to restart a node.
>> >> If you know in advance what you want your cluster to use for, good for
>> >> you, but if somebody wants to do this in runtime why would we force
>> >> them to restart the whole cluster if it is not necessary?
>> >>
>> >> *For that TWCS specifically, what I see is that there might be a TWCS
>> >> table with windows there is only one SSTable in, and if that table is
>> >> e.g. said to be compressed with Zstd while we use flush_compression:
>> >> fast in yaml and it is flushed with LZ4, then it might happen that
>> >> normally (no major compaction and similar, just background compaction)
>> >> that SSTable will not be re-compressed with Zstd. We might have a
>> >> bunch of SSTables on disk compressed with LZ4 while their table
>> >> compression is Zstd just because no compaction covered them (or ruled
>> >> them out). I have a JUnit / bash script simulating this. Maybe this is
>> >> something to fix in TWCS (like re-check that a window with one SSTable
>> >> is on the correct compressor before skipping it) but it is an instance
>> >> where we say to a user that we will compress their tables with one
>> >> compressor but it is compressed with another.
>> >>
>> >> On Mon, Sep 7, 2026 at 3:23 PM Joseph Lynch <[email protected]> wrote:
>> >> >
>> >> > Seems like a reasonable idea to me. I recall that we didn't implement
>> >> > this last time for two reasons: 1) the 4.0 freeze was in effect so we
>> >> > avoided introducing more public surface area (the table option) and 2)
>> >> > we couldn't present a reason at the time why someone would need to set
>> >> > this at a table level. The main case we could think of was if someone
>> >> > was running a backup procedure that (incorrectly) assumed all sstables
>> >> > in Cassandra are compressed identically; this is why we left the
>> >> > "table" option to flip the whole node.
>> >> >
>> >> > Just curious, what is the use case you're seeing where users need to
>> >> > configure this? Is it that they have one Zstd/Deflate table that needs
>> >> > to flush with the same compressor, but they don't want all the others
>> >> > flushing slowly? Maybe a table which defers compaction long enough
>> >> > that the extra disk space before the first deferred compaction becomes
>> >> > significant? I'm just thinking about when we would recommend users
>> >> > deviate from "auto"?
>> >> >
>> >> > Thank you for tackling this!
>> >> >
>> >> > -Joey
>> >> >
>> >> > On Mon, Sep 7, 2026 at 6:52 AM Bernardo Botella
>> >> > <[email protected]> wrote:
>> >> > >
>> >> > > I’m +1 for the functionality.
>> >> > >
>> >> > > From: Štefan Miklošovič <[email protected]>
>> >> > > Date: Monday, 7 September 2026 at 12:07
>> >> > > To: dev <[email protected]>
>> >> > > Subject: [DISCUSS] flush_compression as a parameter in CQL schema
>> >> > >
>> >> > > I want to formally run this (1) through dev ML and gather the approval
>> >> > > for the introduction of this functionality as it technically changes
>> >> > > CQL.
>> >> > >
>> >> > > Thanks and regards
>> >> > >
>> >> > > (1) https://issues.apache.org/jira/browse/CASSANDRA-21642

Reply via email to