PG1204 opened a new issue, #6965:
URL: https://github.com/apache/texera/issues/6965

   ### What happened?
   
   ### What happened?
   
   The HuggingFace inference operator resolves the providers serving a model 
(`_resolve_providers`) and tries them cheapest-first (`_post_with_fallback`). 
Each resolved entry carries a `providerId`, which is the model's name on that 
provider, which provider-scoped router routes 
(`router.huggingface.co/{provider}/...`) require.
   
   For `text-generation` and `image-text-to-text` (the default task path), the 
fallback loop ignores `providerId` and posts `"model": self.MODEL_ID` (the Hub 
ID) to every provider's chat route (`PythonCodegenBase.scala` ~L203-206 in 
`common/workflow-operator/.../huggingFace/codegen/`). Providers whose internal 
model name differs reject the request with 400/404. `_call_provider` (other 
task families) already uses `provider_id` correctly.
   
   The `hf-inference` chat URL is also wrong: built as 
`hf-inference/v1/chat/completions`, but the endpoint requires the model in the 
path (`hf-inference/models/{model-id}/v1/chat/completions`, the form the same 
file already uses for pipeline tasks). Since hf-inference sorts first in 
`PROVIDER_COST_PRIORITY`, the cheapest provider fails on every row.
   
   Effects:
   
   - Text-gen fallback only succeeds when a provider's model name equals the 
Hub ID; otherwise rows fail with "All inference providers failed" despite live 
providers.
   - Every row wastes at least one doomed request (up to the 120 s timeout) 
before falling through, inflating latency and rate-limit pressure.
   
   Fix (contained to the text-gen branch of `_post_with_fallback`): per 
attempt, send a copy of the payload with `"model"` set to that provider's 
`provider_id` (copy, not mutation, the payload is reused across the loop), and 
add the `models/{model-id}` segment to the hf-inference chat URL. Update the 
generated-code assertions in `HuggingFaceInferenceOpDescSpec.scala` to pin both 
behaviors.
   
   ### How to reproduce?
   
   1. Add the operator with task `text-generation` and a model whose 
`providerId` differs from its Hub ID (see 
`https://huggingface.co/api/models/{id}?expand[]=inferenceProviderMapping`).
   2. Run with a valid HF token.
   3. Rows fail or fall back to pricier providers; the generated Python shows 
the chat branch posting `"model": self.MODEL_ID` everywhere and the 
hf-inference chat URL missing `models/{id}`.
   
   ### Version/Branch
   
   1.3.0-incubating-SNAPSHOT (main)
   
   ### Commit Hash (Optional)
   
   _No response_
   
   ### What browsers are you seeing the problem on?
   
   _No response_
   
   ### Relevant log output
   
   ```shell
   
   ```


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to