On Wed, Sep 30, 2026 at 10:27:27PM -0700, David Birdsong wrote: > Concretely: take a multi-tenant backend with, say, 200 application servers, > one hash key per tenant. Today, with map-based or consistent hashing, a > tenant's traffic is deterministic under steady state, but there's no fixed > membership guarantee once servers churn. With consistent hashing > specifically, when a server goes down the ring walk moves that tenant's > traffic to whatever server is next around the ring -- which is a function of > everyone else's current health, not a fixed property of the tenant's key. > Over enough failures or scaling events, a tenant's traffic can, in the > worst case, land on any server in the backend. There's no way to answer, > ahead of time, "which servers can tenant X's requests ever reach" with > anything narrower than "the whole backend, eventually."
You've mentioned this point before but it's not clear to me: what do you call a "multi-tenant backend" ? You seem to imply that you would place multiple applications inside the same backend, but that's strange because a backend currently *is* one instance of an application, with its own rules, load balancing, stickiness etc. That's why I'm not sure that's what you mean and am still confused about your usage of "tenant" here. > That's the problem I want to solve: containment. If a tenant is noisy -- a > bug, an abusive client, a runaway batch job -- its blast radius today is > effectively unbounded. It can, over time, degrade service for every other > tenant sharing the backend, because nothing enforces a ceiling on how many > distinct servers it can touch. It entirely depends on the LB algorithm. If you're using round-robin, of course, all servers will be visited. With leastconn, possibly less but still many. With hashing, it will solely depend on the hashed criteria and how the client can control them (and their willingness to intentionally cause harm of course). For example, consistent-hashing is used a lot with caches because it maintains a high cache hit ratio by sending the same URL to the same server, which means that a client hammering a given URL will not affect other nodes (but the client that scans many URLs will). And combined with the load-factor, it will make use of adjacent nodes for the same URL, spreading the load on the smallest subset needed to handle the load. Other services will hash on the client address to maintain a form of rough stickiness and avoid cache reloads most of the time. > With rendezvous-subset and, say, hash-candidates 3, each tenant key is > ranked against all 200 servers up front, independent of health, and is > permanently bound to its top 3. This is precisely the point I just cannot grasp based on my misunderstanding of what you call a tenant in this context. > Health and load only ever pick among those > 3 (or let us queue/redispatch within them); a tenant literally cannot reach > a 4th server, no matter what else happens in the backend. So "which servers > can tenant X ever reach" becomes a static, computable fact -- three named > servers -- instead of "potentially all of them, given enough churn." That's > the property the ring walk's health-dependent fallback can't give at any > virtual-node count, which is why I think it needs a new algorithm rather > than a tweak to consistent hashing. At this point I got totally lost again :-/ Willy

