The Overseer knows everything.  The clusterstate source of truth *is* the
Overseer's in-memory cache of everything (_not_ ZK).  It writes out its
state changes to ZK but generally doesn't read back changes (perhaps
some?), since it's the gateway for state modification.  I forget if the
Overseer's clusterstate is shared with the node it's running on, and thus
served by whatever requests land there.

The scalability concern is really only when there's a crazy number of
collections; say ~50K+.  Been there ;-)   It's because replicas live off
shards off collections.  It'd be nice if there was a secondary node index,
I suppose.

Any way, the most scalable way to identify the replicas on a specific node
is probably to communicate directly with that node and ask it.  We do have
a "cores" level API; I don't recall if it exposes SolrCloud aspects of the
cores.  And I'm not certain if that would be sufficient for the callers /
use-cases of what's being discussed in this thread, since a core that is
down / non-existent wouldn't show up in /admin/cores.  I am not sure if the
node is watching state for replicas it _should_ have but doesn't yet.  I'm
also unsure if the use-cases of the solr-operator require seeing replica
assignments to nodes when the node is unreachable (was restarted or shut
down indefinitely).

Does the solr-operator or other use cases need to find replicas for a node
that

On Fri, Sep 25, 2026 at 8:29 AM Jan Høydahl <[email protected]> wrote:

> Answering myself.
>
> I initially thought ClusterStateProvider could use such an event API. But
> then it turns out that no single solr node watches all collections in ZK,
> so it would be impossible to implement without watching the world (again).
> Besides, it would be wasteful connection and traffic wise if the client
> only needs on-demand cluster state and not real-time. And I see David
> already plans some new response-haders to drive CSP cache invalidation
> (SOLR-18131).
>
> Jan
>
> > 25. sep. 2026 kl. 00:20 skrev Jan Høydahl <[email protected]>:
> >
> > Hi
> >
> > We could also open the door for thinking very differently about the
> needs of cluster state consumers.
> > Solr nodes themselves subscribe to changes in cluster state from
> Zookeeper through watches.
> > A Solr node does not poll ZK constantly for the entire state tree. But
> we don't want clients to talk directly to ZK.
> >
> > For an app that needs to stay on top of the cluster state (like
> solr-operator),
> > it could be attractive to subscribe to changes as they happen instead of
> polling or on demand pull.
> >
> > In a few projects lately I have implemented SSE (Server Sent Events) in
> backends.
> >
> https://dev.to/abhivyaktii/understanding-server-sent-events-sse-and-why-http2-matters-1cj7
> > It is pure HTTP and easy to implement, understand and to consume.
> > So what if we added an /api/cluster/state/stream endpoint.
> > Then, the server will push new events as it receives updates from ZK
> about nodes coming and going, replicas going down or new collections being
> created.
> > On first (re)connect, server could push the entire cluster state
> initially as a series of different events. Connection is long lived.
> >
> > A stream contains one or more event types, with arbitrary text. We
> choose granularity of events, below is an example:
> >
> > HTTP/1.1 200 OK
> > Content-Type: text/event-stream
> > Cache-Control: no-cache
> > Connection: keep-alive
> >
> > livenodes: ["127.0.1.1:7574_solr", "127.0.1.1:7500_solr",
> "127.0.1.1:8983_solr", "127.0.1.1:8900_solr"]
> >
> > aliases: {"both_collections":"collection1,collection2"}
> >
> > roles: {"overseer": ["127.0.1.1:8983_solr", "127.0.1.1:7574_solr"]}
> >
> > collection: {"collection1":{"shards": {"shard1":{"state":"active"}}}}
> >
> > collection: {"collection2":{"shards": {"shard1":{"state":"down"}}}}
> >
> > livenodes: ["127.0.1.1:7574_solr", "127.0.1.1:7500_solr",
> "127.0.1.1:8983_solr"]  // node down
> >
> > collection: {"collection2":{}}   // deletion
> >
> >
> > Clients that do not need to keep track of the entire state can simply
> request bits and pieces on demand from the other endpoints.
> >
> > Jan
> >
> >
> >> 24. sep. 2026 kl. 15:39 skrev Jason Gerlowski <[email protected]>:
> >>
> >> Hey Christos,
> >>
> >> Do you have a particular division or set of endpoints in mind that
> >> you'd recommend?
> >>
> >> The tension I'm struggling with here is the difference between the
> >> design that's the most "REST-ful" and the one that's the most useful.
> >> Those seem to be at odds here.  CLUSTERSTATUS is an ugly,
> >> mega-endpoint that returns way too much information....but if you look
> >> at its callers they tend to want/need all that information.  They're
> >> making bounce or replica placement decisions that require knowing the
> >> status of...well, the whole cluster.  We can decompose the endpoint,
> >> but if the result is that callers would now need to make 10 or 20 or
> >> 100 API calls (as in epugh's example above) it's a worse experience
> >> for them.
> >>
> >> It seems like you're suggesting that conflict is avoidable with the
> >> right resource definitions or endpoints, but I'm having trouble seeing
> >> what you have in mind.  Could you give a bit more detail please?
> >>
> >> Best,
> >>
> >> Jason
> >>
> >> On Tue, Sep 22, 2026 at 7:37 AM Christos Malliaridis
> >> <[email protected]> wrote:
> >>>
> >>> Sorry for entering the discussion so late with my two cents below, but
> perhaps I can add some more insights from a consumer's perspective and help
> answer the question "to what direction should we head with v2 API".
> >>>
> >>> In a previous discussion a couple months ago, Jason and I were talking
> about "REST"ful APIs, mainly from a consumer's perspective, but also for
> determining the API endpoints. We were trying to define what the structure
> of the API endpoints should be, and what data should be provided by each
> endpoint, with the goal to also avoid duplicated data across endpoints.
> >>>
> >>> In a RESTful API, we are normally talking about resources. The v2
> proposal page in the spreadsheet [1] was aiming for defining the resources
> we have in Solr at API level, without being influenced from internal
> structure or implementation details.
> >>>
> >>> Taking the discussion's outcome, /api/cluster would fetch a single
> resource element, the cluster, and provide information about the cluster.
> Not the nodes, not the shards, not collections. If we are interested in
> nodes, we would use the resource collection GET /api/nodes, optionally with
> some query parameters like status=healthy.
> >>>
> >>> By focusing on what a resource is and what information it carries
> (scoped), we would avoid super-endpoints that provide too much information,
> like the CLUSTERSTATUS does right now.
> >>>
> >>> The new Admin UI is also applying the concept of resources in the
> designs. I am not sure if we want to follow that ideology, but it could
> definitely help against excessive data exposure from single endpoints. The
> SOLID principles would also take effect in various ways.
> >>>
> >>> [1]
> https://docs.google.com/spreadsheets/d/1HAoBBFPpSiT8mJmgNZKkZAPwfCfPvlc08m5jz3fQBpA/edit?pli=1&gid=1878317994#gid=1878317994
> >>>
> >>> On 2026/09/08 11:19:26 Eric Pugh wrote:
> >>>> Hi all, spelunking a bit through what API to convert from our old
> homegrown V2 and to the Jax RS V2 approach, and I noticed that
> CLUSTERSTATUS is one that is used to drive a lot of our Admin UI.   So I
> thought, hey, that will be an easy migration.
> >>>>
> >>>> I found this closed (but not merged) PR:
> https://github.com/apache/solr/pull/2670 from David, and this JIRA
> https://issues.apache.org/jira/browse/SOLR-17422.
> >>>>
> >>>> Before I go down the path of making some more JIRAs for various V2
> api work for folks to pick up, at least for CLUSTERSTATUS, I wanted to see
> if anyone had sketched out what we WANT the API structure to look like?
> >>>>
> >>>> I had a comment from two years ago on this, but hadn’t done anything
> about it…..
> >>>>
> >>>> Eric
> >>>>
> >>>> Disclaimer
> >>>>
> >>>> The information contained in this communication from the sender is
> confidential. It is intended solely for use by the recipient and others
> authorized to receive it. If you are not the recipient, you are hereby
> notified that any disclosure, copying, distribution or taking action in
> relation of the contents of this information is strictly prohibited and may
> be unlawful.
> >>>>
> >>>> This email has been scanned for viruses and malware, and may have
> been automatically archived by Mimecast, a leader in email security and
> cyber resilience. Mimecast integrates email defenses with brand protection,
> security awareness training, web security, compliance and other essential
> capabilities. Mimecast helps protect large and small organizations from
> malicious activity, human error and technology failure; and to lead the
> movement toward building a more resilient world. To find out more, visit
> our website.
> >>>>
> >>>
> >>> ---------------------------------------------------------------------
> >>> To unsubscribe, e-mail: [email protected]
> >>> For additional commands, e-mail: [email protected]
> >>>
> >>
> >> ---------------------------------------------------------------------
> >> To unsubscribe, e-mail: [email protected]
> >> For additional commands, e-mail: [email protected]
> >>
> >
>
>

Reply via email to