Hi Andrew, >Is it really useful to expose this concept to normal applications? Yes, it's required to have pod in metadata in order to extend the concept of isolation from replica-broker placement to producer/consumer-broker traffic routing. for example, blast radius of a bad kafka client deployment e.g. connection leak can be constrained within canary producers. concretely client can use pod in the metadata to select partitions and achieve isolation. (client pod isolation is a future improvement not completed defined in the KIP, change of MetadataResponse is for future extension)
>There is a multi-tenancy KIP and I imagine that we would want to be able to scope tenants to a subset of brokers. I'm aware of the KIP. but physical isolation by tenant is still pretty ambiguous for kafka. not like storage/database, pub/sub naturally involves processor of different service/org/tenant. normally topic producer is one tenant, but consumer could be another >It would be much preferable if the tool could discover the canary information from the cluster to ensure consistency Make sense, will improve > I suggest "cell" instead, but this is just my personal opinion. Great suggestion, cell should be a better name. considering it's a adopted term in cloud providers such as GCP and AWS. will change later if there is no other better alternative term. Thanks Andrew Schofield <[email protected]> 于2026年8月9日周日 01:18写道: > Hi Zhifeng, > Thanks for reviving this KIP. I'm very supportive of the principle of > being able to subdivide the brokers in a cluster for a variety of reasons. > I have a few initial comments. > > AS1: I do not think the pod should be exposed in the MetadataResponse and > o.a.k.common.Node. Is it really useful to expose this concept to normal > applications? The reason I am asking is that I think we might want to > subdivide clusters in a way which is invisible to applications in the > future. There is a multi-tenancy KIP and I imagine that we would want to be > able to scope tenants to a subset of brokers. > > AS2: It seems a bit inelegant having the canary configuration in the > controller configs and also in the arguments to > kafka-reassign-partitions.sh. It would be much preferable if the tool could > discover the canary information from the cluster to ensure consistency. > > AS3: I wonder whether using "pod" when it is so widely used in the context > of Kubernetes is sensible. I suggest "cell" instead, but this is just my > personal opinion. > > Thanks, > Andrew > > On 2026/08/06 07:02:27 Chen Zhifeng wrote: > > Title: [DISCUSS] KIP-1095 Kafka Canary Isolation > > > > Hi Everyone, > > > > Apologies for the long silence. Reviving KIP-1095 > > < > https://cwiki.apache.org/confluence/spaces/KAFKA/pages/323488210/KIP-1095+Kafka+Canary+Isolation > > > > after a substantial revision. The original thread is at here > > <https://lists.apache.org/thread/n7mprq43fh39hsgj48bzfbtbho3l1cpy>; and > > thanks Divij for the questions there. > > > > *What changed since initial discussion* > > - Scope narrowed to broker-side placement - producer/consumer-side > > improvements as future work > > - Cleaned up zookeeper related changes > > - Protocol changes reframed as tagged fields - no client upgrade > required > > > > *Answering the earlier questions* > > > Why can’t we achieve the objective without making any change at all? > For > > example, you can designate a few brokers as your “canary brokers” where > > your custom "canary partitions" are situated. During rolling deployment > you > > can choose to deploy changes to these brokers at the beginning. If the > > health of your canary partitions is good, you can continue ahead with the > > rest of deployment. > > A: The suggested deployment process has been used at Uber for years, and > > this KIP is developed on top of it. What operating it showed is that > > designating brokers alone does not give you isolation between the > > designated brokers and the rest. > > A partition's replicas span brokers, so unless placement guarantees that > > the entire replica set lands inside the canary pod, a canary partition > > still has followers on non-canary brokers, and vice versa. Without that > > isolation, impact leaks from one broker to every other broker it is > > connected to — a leader running new code can propagate bad data to > > followers, and a degraded follower can slow down a leader that was never > > upgraded. The blast radius of the deployment therefore grows from a > small, > > known percentage of requests to an unknown and unbounded portion. > > Detection suffers for the same reason. When one partition degrades, > > producers of keyless records simply route around it to healthy > partitions, > > so a partial failure may not be detectable until the rollout is nearly > > complete. > > > > > What do you think about having a separate canary cluster where you > deploy > > code first before deploying to production cluster. The canary cluster > could > > receive a small portion of "shadow" production traffic or have it's own > > synthetic traffic. > > A: Shadowing, or capture/replay, is another approach to safe deployment, > > with a different trade-off: > > - capture/replay pays for redundancy in order to keep impact away from > > production entirely, which requires extra storage, compute, and > engineering > > cost; > > - canary isolation instead bounds the blast radius of a bad deployment > to > > a pre-calculated portion of traffic, at lower engineering cost and with > no > > extra hardware. > > One important difference is that shadow traffic does not reproduce > > production behaviour exactly, so some regressions will always slip > through > > to production. Canary traffic is production traffic, so it inherits > > production behaviour by construction. The two are complementary, and > which > > one fits depends on the operator; this KIP aims to make the second option > > available in Kafka itself. > > > > > Would controller broker be part of canary brokers or not? How will we > > test code regression in controller? Similarly how will we test code > > regression in transaction coordinator and consumer coordinator? > > A: controllers are not covered by this KIP. In KRaft the controller is a > > separate role with its own quorum and rolling procedure, and the > > canary-partition concept does not map onto it; catching controller > > regressions needs a different mechanism, which I would rather not fold > into > > this proposal. > > > > The coordinators are a different case. Both the group and transaction > > coordinators are partition leaders of __consumer_offsets and > > __transaction_state, so the mechanism in this KIP reaches them by > applying > > the same placement rules to those internal topics. I have left that out > of > > the initial scope deliberately, but I am happy to discuss it as follow-up > > work if there is interest. > > > > *Production status*: The idea has been applied at Uber for around 2 > years. > > with 1/32 partitions being canary partitions, blast radius of bad kafka > > deployment has been contained with-in ~3% of production. > > > > *Draft implementation*: https://github.com/apache/kafka/pull/23095. > > > > *JIRA*: https://issues.apache.org/jira/browse/KAFKA-20897. > > > > Thanks, > > Zhifeng > > >
