The Overseer knows everything. The clusterstate source of truth *is* the Overseer's in-memory cache of everything (_not_ ZK). It writes out its state changes to ZK but generally doesn't read back changes (perhaps some?), since it's the gateway for state modification. I forget if the Overseer's clusterstate is shared with the node it's running on, and thus served by whatever requests land there.
The scalability concern is really only when there's a crazy number of collections; say ~50K+. Been there ;-) It's because replicas live off shards off collections. It'd be nice if there was a secondary node index, I suppose. Any way, the most scalable way to identify the replicas on a specific node is probably to communicate directly with that node and ask it. We do have a "cores" level API; I don't recall if it exposes SolrCloud aspects of the cores. And I'm not certain if that would be sufficient for the callers / use-cases of what's being discussed in this thread, since a core that is down / non-existent wouldn't show up in /admin/cores. I am not sure if the node is watching state for replicas it _should_ have but doesn't yet. I'm also unsure if the use-cases of the solr-operator require seeing replica assignments to nodes when the node is unreachable (was restarted or shut down indefinitely). Does the solr-operator or other use cases need to find replicas for a node that On Fri, Sep 25, 2026 at 8:29 AM Jan Høydahl <[email protected]> wrote: > Answering myself. > > I initially thought ClusterStateProvider could use such an event API. But > then it turns out that no single solr node watches all collections in ZK, > so it would be impossible to implement without watching the world (again). > Besides, it would be wasteful connection and traffic wise if the client > only needs on-demand cluster state and not real-time. And I see David > already plans some new response-haders to drive CSP cache invalidation > (SOLR-18131). > > Jan > > > 25. sep. 2026 kl. 00:20 skrev Jan Høydahl <[email protected]>: > > > > Hi > > > > We could also open the door for thinking very differently about the > needs of cluster state consumers. > > Solr nodes themselves subscribe to changes in cluster state from > Zookeeper through watches. > > A Solr node does not poll ZK constantly for the entire state tree. But > we don't want clients to talk directly to ZK. > > > > For an app that needs to stay on top of the cluster state (like > solr-operator), > > it could be attractive to subscribe to changes as they happen instead of > polling or on demand pull. > > > > In a few projects lately I have implemented SSE (Server Sent Events) in > backends. > > > https://dev.to/abhivyaktii/understanding-server-sent-events-sse-and-why-http2-matters-1cj7 > > It is pure HTTP and easy to implement, understand and to consume. > > So what if we added an /api/cluster/state/stream endpoint. > > Then, the server will push new events as it receives updates from ZK > about nodes coming and going, replicas going down or new collections being > created. > > On first (re)connect, server could push the entire cluster state > initially as a series of different events. Connection is long lived. > > > > A stream contains one or more event types, with arbitrary text. We > choose granularity of events, below is an example: > > > > HTTP/1.1 200 OK > > Content-Type: text/event-stream > > Cache-Control: no-cache > > Connection: keep-alive > > > > livenodes: ["127.0.1.1:7574_solr", "127.0.1.1:7500_solr", > "127.0.1.1:8983_solr", "127.0.1.1:8900_solr"] > > > > aliases: {"both_collections":"collection1,collection2"} > > > > roles: {"overseer": ["127.0.1.1:8983_solr", "127.0.1.1:7574_solr"]} > > > > collection: {"collection1":{"shards": {"shard1":{"state":"active"}}}} > > > > collection: {"collection2":{"shards": {"shard1":{"state":"down"}}}} > > > > livenodes: ["127.0.1.1:7574_solr", "127.0.1.1:7500_solr", > "127.0.1.1:8983_solr"] // node down > > > > collection: {"collection2":{}} // deletion > > > > > > Clients that do not need to keep track of the entire state can simply > request bits and pieces on demand from the other endpoints. > > > > Jan > > > > > >> 24. sep. 2026 kl. 15:39 skrev Jason Gerlowski <[email protected]>: > >> > >> Hey Christos, > >> > >> Do you have a particular division or set of endpoints in mind that > >> you'd recommend? > >> > >> The tension I'm struggling with here is the difference between the > >> design that's the most "REST-ful" and the one that's the most useful. > >> Those seem to be at odds here. CLUSTERSTATUS is an ugly, > >> mega-endpoint that returns way too much information....but if you look > >> at its callers they tend to want/need all that information. They're > >> making bounce or replica placement decisions that require knowing the > >> status of...well, the whole cluster. We can decompose the endpoint, > >> but if the result is that callers would now need to make 10 or 20 or > >> 100 API calls (as in epugh's example above) it's a worse experience > >> for them. > >> > >> It seems like you're suggesting that conflict is avoidable with the > >> right resource definitions or endpoints, but I'm having trouble seeing > >> what you have in mind. Could you give a bit more detail please? > >> > >> Best, > >> > >> Jason > >> > >> On Tue, Sep 22, 2026 at 7:37 AM Christos Malliaridis > >> <[email protected]> wrote: > >>> > >>> Sorry for entering the discussion so late with my two cents below, but > perhaps I can add some more insights from a consumer's perspective and help > answer the question "to what direction should we head with v2 API". > >>> > >>> In a previous discussion a couple months ago, Jason and I were talking > about "REST"ful APIs, mainly from a consumer's perspective, but also for > determining the API endpoints. We were trying to define what the structure > of the API endpoints should be, and what data should be provided by each > endpoint, with the goal to also avoid duplicated data across endpoints. > >>> > >>> In a RESTful API, we are normally talking about resources. The v2 > proposal page in the spreadsheet [1] was aiming for defining the resources > we have in Solr at API level, without being influenced from internal > structure or implementation details. > >>> > >>> Taking the discussion's outcome, /api/cluster would fetch a single > resource element, the cluster, and provide information about the cluster. > Not the nodes, not the shards, not collections. If we are interested in > nodes, we would use the resource collection GET /api/nodes, optionally with > some query parameters like status=healthy. > >>> > >>> By focusing on what a resource is and what information it carries > (scoped), we would avoid super-endpoints that provide too much information, > like the CLUSTERSTATUS does right now. > >>> > >>> The new Admin UI is also applying the concept of resources in the > designs. I am not sure if we want to follow that ideology, but it could > definitely help against excessive data exposure from single endpoints. The > SOLID principles would also take effect in various ways. > >>> > >>> [1] > https://docs.google.com/spreadsheets/d/1HAoBBFPpSiT8mJmgNZKkZAPwfCfPvlc08m5jz3fQBpA/edit?pli=1&gid=1878317994#gid=1878317994 > >>> > >>> On 2026/09/08 11:19:26 Eric Pugh wrote: > >>>> Hi all, spelunking a bit through what API to convert from our old > homegrown V2 and to the Jax RS V2 approach, and I noticed that > CLUSTERSTATUS is one that is used to drive a lot of our Admin UI. So I > thought, hey, that will be an easy migration. > >>>> > >>>> I found this closed (but not merged) PR: > https://github.com/apache/solr/pull/2670 from David, and this JIRA > https://issues.apache.org/jira/browse/SOLR-17422. > >>>> > >>>> Before I go down the path of making some more JIRAs for various V2 > api work for folks to pick up, at least for CLUSTERSTATUS, I wanted to see > if anyone had sketched out what we WANT the API structure to look like? > >>>> > >>>> I had a comment from two years ago on this, but hadn’t done anything > about it….. > >>>> > >>>> Eric > >>>> > >>>> Disclaimer > >>>> > >>>> The information contained in this communication from the sender is > confidential. It is intended solely for use by the recipient and others > authorized to receive it. If you are not the recipient, you are hereby > notified that any disclosure, copying, distribution or taking action in > relation of the contents of this information is strictly prohibited and may > be unlawful. > >>>> > >>>> This email has been scanned for viruses and malware, and may have > been automatically archived by Mimecast, a leader in email security and > cyber resilience. Mimecast integrates email defenses with brand protection, > security awareness training, web security, compliance and other essential > capabilities. Mimecast helps protect large and small organizations from > malicious activity, human error and technology failure; and to lead the > movement toward building a more resilient world. To find out more, visit > our website. > >>>> > >>> > >>> --------------------------------------------------------------------- > >>> To unsubscribe, e-mail: [email protected] > >>> For additional commands, e-mail: [email protected] > >>> > >> > >> --------------------------------------------------------------------- > >> To unsubscribe, e-mail: [email protected] > >> For additional commands, e-mail: [email protected] > >> > > > >
