This is an automated email from the ASF dual-hosted git repository.

github-merge-queue[bot] pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/texera.git


The following commit(s) were added to refs/heads/main by this push:
     new 378a3b54ec feat(workflow-operator): read chat-provider responses for 
the image question-answering tasks (#7920)
378a3b54ec is described below

commit 378a3b54ec645e57760011bca0e08c1829c86680
Author: Prateek Ganigi <[email protected]>
AuthorDate: Thu Sep 17 22:16:29 2026 +0000

    feat(workflow-operator): read chat-provider responses for the image 
question-answering tasks (#7920)
    
    ### What changes were proposed in this PR?
    
    When the operator falls back from `hf-inference` to a third-party
    chat-completions provider, the reply comes back as `{"choices":
    [{"message": {"content": ...}}]}`. Three image tasks in
    `ImageTaskCodegen.parsePython` could not read that shape, so a correct
    answer was written to the result column as a raw JSON envelope:
    
    - `visual-question-answering` and `document-question-answering` returned
    `body.get("answer", json.dumps(body))`, and a chat response has no
    `answer` key.
    - `zero-shot-image-classification` shared the image-only branch, which
    always returns`json.dumps(body)`.
    
    Both now read `choices[0]["message"]["content"]` when the body carries
    `choices`, keeping the native `hf-inference` shape as the primary path.
    `zero-shot-image-classification` gets its own branch, placed ahead of
    the image-only tasks because the generated `if/elif` chain is
    first-match-wins. This is the same idiom `image-to-text` and
    `image-text-to-text` already use in this file, and the one applied to
    the text tasks in #7798.
    
    `image-classification`, `object-detection` and `image-segmentation` are
    left as they are: they have no question to answer, so a free-text chat
    reply is not meaningful structured output for them.
    
    This is Part A of #7906 and covers the response side only. The request
    side which is carrying `candidate_labels` into the chat message for
    `zero-shot-image-classification`, follows in Part B.
    
    ### Any related issues?
    
    Addresses #7906
    
    ### How was this PR tested?
    
    133 tests pass in the `WorkflowOperator` Hugging Face suites,
    `PythonCodeRawInvalidTextSpec` py-compiles the generated Python for all
    117 operators, and `scalafmtCheck` is clean for main and test sources.
    Two tests were added to `ImageTaskCodegenSpec`: one asserts the
    visual/document question-answering branch reads `choices` ahead of the
    native `answer` lookup, the other asserts the new
    `zero-shot-image-classification` branch exists and precedes the
    image-only branch.
    
    The emitted Python was also exercised directly: the three fixed tasks
    return the chat content, native `hf-inference` responses parse exactly
    as before, non-dict and `answer`-less bodies still fall through to
    `json.dumps`, and the untouched branches (`image-classification`,
    `object-detection`, `image-segmentation`, `image-to-text`,
    `image-text-to-text`) are unchanged.
    
    ### Was this PR authored or co-authored using generative AI tooling?
    
    Yes, this PR was co-authored with Claude in compliance with ASF policy.
    
    ---------
    
    Co-authored-by: Xuan Gu <[email protected]>
---
 .../huggingFace/codegen/ImageTaskCodegen.scala     | 18 ++++++---
 .../huggingFace/codegen/ImageTaskCodegenSpec.scala | 47 ++++++++++++++++++++++
 2 files changed, 60 insertions(+), 5 deletions(-)

diff --git 
a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
 
b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
index 673227ab9e..201d7b1623 100644
--- 
a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
+++ 
b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
@@ -106,18 +106,22 @@ object ImageTaskCodegen extends TaskCodegen {
       |                if isinstance(body, dict):
       |                    if "md_results" in body:
       |                        return body["md_results"]
-      |                    if "choices" in body:
-      |                        return body["choices"][0]["message"]["content"]
+      |                    if body.get("choices"):
+      |                        return body["choices"][0].get("message", 
{}).get("content", json.dumps(body))
       |                if isinstance(body, list) and body and 
isinstance(body[0], dict):
       |                    return body[0].get("generated_text", 
json.dumps(body))
       |                return json.dumps(body)
       |            elif task in ("visual-question-answering", 
"document-question-answering"):
       |                if isinstance(body, dict):
+      |                    # Third-party chat providers answer via 
choices[0].message;
+      |                    # hf-inference returns the native {"answer": ...} 
shape.
+      |                    if body.get("choices"):
+      |                        return body["choices"][0].get("message", 
{}).get("content", json.dumps(body))
       |                    return body.get("answer", json.dumps(body))
       |                return json.dumps(body)
       |            elif task == "image-text-to-text":
-      |                if isinstance(body, dict) and "choices" in body:
-      |                    return body["choices"][0]["message"]["content"]
+      |                if isinstance(body, dict) and body.get("choices"):
+      |                    return body["choices"][0].get("message", 
{}).get("content", json.dumps(body))
       |                if isinstance(body, list) and body and 
isinstance(body[0], dict):
       |                    return body[0].get("generated_text", 
json.dumps(body))
       |                return json.dumps(body)
@@ -144,6 +148,10 @@ object ImageTaskCodegen extends TaskCodegen {
       |                            if "url" in data[0]:
       |                                return 
self._url_to_data_url(data[0]["url"])
       |                return json.dumps(body)
-      |            elif task in ("image-classification", "object-detection", 
"image-segmentation", "zero-shot-image-classification"):
+      |            elif task == "zero-shot-image-classification":
+      |                if isinstance(body, dict) and body.get("choices"):
+      |                    return body["choices"][0].get("message", 
{}).get("content", json.dumps(body))
+      |                return json.dumps(body)
+      |            elif task in ("image-classification", "object-detection", 
"image-segmentation"):
       |                return json.dumps(body)""".stripMargin
 }
diff --git 
a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
 
b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
index f1806d3b94..0e026fdf0f 100644
--- 
a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
+++ 
b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
@@ -116,6 +116,53 @@ class ImageTaskCodegenSpec extends AnyFlatSpec with 
Matchers {
     out should include("json.dumps(body)")
   }
 
+  it should "read chat-provider responses for visual and document 
question-answering" in {
+    // #7906 Part A: these tasks only understood hf-inference's native
+    // {"answer": ...} shape, so a chat-completions reply was written to the
+    // result column as a raw JSON envelope. The native shape stays primary.
+    val out = ImageTaskCodegen.parsePython(makeCtx())
+    val branch = out
+      .split("""elif task """)
+      .find(_.startsWith("""in ("visual-question-answering"""))
+      .getOrElse(fail("visual/document question-answering branch is missing"))
+    branch should include(
+      """body["choices"][0].get("message", {}).get("content", 
json.dumps(body))"""
+    )
+    branch should include("""body.get("answer"""")
+    branch.indexOf("choices") should be < 
branch.indexOf("""body.get("answer"""")
+  }
+
+  it should "read chat-provider responses for zero-shot-image-classification" 
in {
+    // #7906 Part A: the task used to share the image-only branch, which always
+    // dumps the body, so a chat reply was never extracted. It now has its own
+    // branch, which must precede the image-only one because the generated
+    // if/elif chain is first-match-wins.
+    val out = ImageTaskCodegen.parsePython(makeCtx())
+    out should include("""elif task == "zero-shot-image-classification":""")
+    out should include(
+      """elif task in ("image-classification", "object-detection", 
"image-segmentation"):"""
+    )
+    out.indexOf("""elif task == "zero-shot-image-classification":""") should 
be <
+      out.indexOf("""elif task in ("image-classification",""")
+  }
+
+  it should "degrade instead of raising when a chat response is malformed" in {
+    // Review feedback on #7920: parsePython runs per row, so indexing straight
+    // into choices[0]["message"]["content"] turns one malformed provider
+    // response into an aborted run — an empty "choices" list raises IndexError
+    // and a choice without "message"/"content" raises KeyError. Every chat
+    // extraction in this file now uses a truthiness guard plus .get chaining,
+    // matching how the native shapes already degrade via .get(..., 
json.dumps(body)).
+    val out = ImageTaskCodegen.parsePython(makeCtx())
+    out should not include ("""["message"]["content"]""")
+    out.split("""body\.get\("choices"\)""").length - 1 shouldBe 4
+    out
+      .split(
+        """\.get\("message", \{\}\)\.get\("content", json\.dumps\(body\)\)"""
+      )
+      .length - 1 shouldBe 4
+  }
+
   "ImageTaskCodegen snippets" should "never inline raw CodegenContext string 
values" in {
     // The snippets are static and reference only self.* attributes; the base
     // class decodes user-supplied strings safely at runtime. Sentinel values

Reply via email to