ashoka1981 opened a new issue, #68626:
URL: https://github.com/apache/doris/issues/68626

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/doris/issues?q=is%3Aissue) and found no 
similar issues.
   
   ### Description
   
   The scalar AI functions (`ai_filter`, `ai_classify`, `ai_sentiment`, 
`ai_summarize`, ... — everything built on `AIFunction<Derived>::execute` in 
`be/src/exprs/function/ai/ai_functions.h`) batch rows into one provider request 
per `ai_context_window_size` bytes (since #62494), but the batches of a block 
are issued strictly one after another: build a batch, block on the HTTP round 
trip, append results, build the next batch.
   
   For the common shape `SELECT id, col, ai_filter('res', ...) FROM t` the 
planner produces a single fragment whose scan projection feeds the result sink 
directly (one instance — `Total Instances Num: 1` in the profile), so the whole 
query has exactly one provider request in flight at any time regardless of 
`parallel_pipeline_task_num`, tablet count or cores. Query time is the sum of 
every batch's provider latency.
   
   Proposal: add an AI resource property `ai.max_concurrency` (default `1` = 
today's behaviour) that lets one execution instance keep up to N batch requests 
in flight, executed on a bounded BE-wide thread pool with a streaming scheduler 
(refill on each completion), results still appended in batch order so output is 
unchanged. The limit lives on the resource because it is a property of the 
provider endpoint (rate limit / connection budget), not of a query, and can be 
changed with `ALTER RESOURCE`.
   
   Measured on a single-BE cluster, 20,000 rows, `ai_filter` with GPT-4.1-mini 
via Azure OpenAI (22 rows/batch): 2,027 s serial → 491 s with 
`ai.max_concurrency=4`, provider latency unchanged, identical results. Against 
a low-latency local provider, 100,000 rows went from 96 s to 50 s.
   
   ### Use case
   
   Any AI function over more than a few thousand rows against a hosted model: 
today the BE is idle waiting on one HTTP call at a time, so throughput is 
bounded by provider latency instead of the provider's concurrency limit. Users 
with a provider that allows parallel requests (all hosted APIs do) get a 3–5x 
speedup by setting one property.
   
   ### Related issues
   
   #62494 introduced the per-block batching this builds on.
   
   ### Are you willing to submit PR?
   
   - [x] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to