[
https://issues.apache.org/jira/browse/SOLR-18466?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Mathew updated SOLR-18466:
--------------------------
Summary: Support _route_ in the streaming expression search function to
prune shard fan-out (was: Support _route_ in the streaming expression function
to prune shard fan-out)
> Support _route_ in the streaming expression search function to prune shard
> fan-out
> ----------------------------------------------------------------------------------
>
> Key: SOLR-18466
> URL: https://issues.apache.org/jira/browse/SOLR-18466
> Project: Solr
> Issue Type: Improvement
> Components: SolrJ, streaming expressions
> Reporter: Mathew
> Priority: Minor
>
> The \{{search()}} function in streaming expressions currently ignores
> \{{_route_}} for shard selection and queries every active slice of the target
> collection. Users with known routing keys therefore incur unnecessary request
> and response overhead, which can be significant on large clusters. This
> change adds support for \{{_route_}} param in the streaming expression search
> function, reusing \{{docCol.getRouter().getSearchSlices()}} which is used for
> shard selection in normal distributed search.
> *Changes:*
> Two files in \{{solr/solrj-streaming}}:
> *1. the CloudSolrStream class:*
> the existing method:
> {code:java}
> List<Slice> getSlices(
> String collectionName,
> CloudSolrClient cloudSolrClient,
> boolean checkAlias)
> {code}
> is kept for existing calls from \{{TopicStream}}, \{{TextLogitStream}},
> \{{FeaturesSelectionStream}}, and \{{StatementImpl}} (all out of scope for
> this change); these are unaffected and continue to receive all active slices
> A new overloaded helper method is added:
> {code:java}
> List<Slice> getSlices(
> String collectionName,
> CloudSolrClient cloudSolrClient,
> boolean checkAlias,
> SolrParams solrParams)
> {code}
> When \{{solrParams}} contains a nonempty \{{_route_}}, each collection's
> slices are resolved via \{{docCol.getRouter().getSearchSlices(routeKeys,
> params, docCol)}} and when \{{_route_}} is absent (or empty), all active
> slices are returned exactly as before.
> *2. the TupleStream class:*
> {\{CloudSolrStream.getSlices(collection, cloudSolrClient, true, solrParams)}}
> is called after the request params are assembled and passed to the route
> aware \{{getSlices}}.
> *Example usages:*
> {noformat}
> expr=search(collection1,
> zkHost="localhost:9983",
> qt="/export",
> q="*:*",
> fl="id,a_s,a_i,a_f",
> sort="a_f asc, a_i asc", _route_="routeKeyA")
> {noformat}
> {noformat}
> expr=search(collection1,
> zkHost="localhost:9983",
> qt="/export",
> q="*:*",
> fl="id,a_s,a_i,a_f",
> sort="a_f asc, a_i asc", _route_="routeKeyA,routeKeyB,routeKeyC")
> {noformat}
> *One note on backward compatibility:*
> Previously, \{{_route_}} was ignored during shard selection for streaming
> expression \{{search()}}. With this change, nonempty route values are passed
> to \{{docCol.getRouter().getSearchSlices(routeKeys, params, docCol)}}. For
> collections using the implicit router, an invalid \{{_route_}} value now
> causes the request to fail. Previously, the same request searched all active
> shards despite the invalid \{{_route_}}. This should be documented in the
> release/ upgrade notes.
> Tests cover single and multiple route keys, composite id routing, missing and
> empty route values, and invalid implicit router shard names.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]