This is an automated email from the ASF dual-hosted git repository. github-merge-queue[bot] pushed a commit to branch gh-readonly-queue/main/pr-7920-471e53cfc1d5f96748a7ec6b7563561239229e5e in repository https://gitbox.apache.org/repos/asf/texera.git
commit 378a3b54ec645e57760011bca0e08c1829c86680 Author: Prateek Ganigi <[email protected]> AuthorDate: Thu Sep 17 22:16:29 2026 +0000 feat(workflow-operator): read chat-provider responses for the image question-answering tasks (#7920) ### What changes were proposed in this PR? When the operator falls back from `hf-inference` to a third-party chat-completions provider, the reply comes back as `{"choices": [{"message": {"content": ...}}]}`. Three image tasks in `ImageTaskCodegen.parsePython` could not read that shape, so a correct answer was written to the result column as a raw JSON envelope: - `visual-question-answering` and `document-question-answering` returned `body.get("answer", json.dumps(body))`, and a chat response has no `answer` key. - `zero-shot-image-classification` shared the image-only branch, which always returns`json.dumps(body)`. Both now read `choices[0]["message"]["content"]` when the body carries `choices`, keeping the native `hf-inference` shape as the primary path. `zero-shot-image-classification` gets its own branch, placed ahead of the image-only tasks because the generated `if/elif` chain is first-match-wins. This is the same idiom `image-to-text` and `image-text-to-text` already use in this file, and the one applied to the text tasks in #7798. `image-classification`, `object-detection` and `image-segmentation` are left as they are: they have no question to answer, so a free-text chat reply is not meaningful structured output for them. This is Part A of #7906 and covers the response side only. The request side which is carrying `candidate_labels` into the chat message for `zero-shot-image-classification`, follows in Part B. ### Any related issues? Addresses #7906 ### How was this PR tested? 133 tests pass in the `WorkflowOperator` Hugging Face suites, `PythonCodeRawInvalidTextSpec` py-compiles the generated Python for all 117 operators, and `scalafmtCheck` is clean for main and test sources. Two tests were added to `ImageTaskCodegenSpec`: one asserts the visual/document question-answering branch reads `choices` ahead of the native `answer` lookup, the other asserts the new `zero-shot-image-classification` branch exists and precedes the image-only branch. The emitted Python was also exercised directly: the three fixed tasks return the chat content, native `hf-inference` responses parse exactly as before, non-dict and `answer`-less bodies still fall through to `json.dumps`, and the untouched branches (`image-classification`, `object-detection`, `image-segmentation`, `image-to-text`, `image-text-to-text`) are unchanged. ### Was this PR authored or co-authored using generative AI tooling? Yes, this PR was co-authored with Claude in compliance with ASF policy. --------- Co-authored-by: Xuan Gu <[email protected]> --- .../huggingFace/codegen/ImageTaskCodegen.scala | 18 ++++++--- .../huggingFace/codegen/ImageTaskCodegenSpec.scala | 47 ++++++++++++++++++++++ 2 files changed, 60 insertions(+), 5 deletions(-) diff --git a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala index 673227ab9e..201d7b1623 100644 --- a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala +++ b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala @@ -106,18 +106,22 @@ object ImageTaskCodegen extends TaskCodegen { | if isinstance(body, dict): | if "md_results" in body: | return body["md_results"] - | if "choices" in body: - | return body["choices"][0]["message"]["content"] + | if body.get("choices"): + | return body["choices"][0].get("message", {}).get("content", json.dumps(body)) | if isinstance(body, list) and body and isinstance(body[0], dict): | return body[0].get("generated_text", json.dumps(body)) | return json.dumps(body) | elif task in ("visual-question-answering", "document-question-answering"): | if isinstance(body, dict): + | # Third-party chat providers answer via choices[0].message; + | # hf-inference returns the native {"answer": ...} shape. + | if body.get("choices"): + | return body["choices"][0].get("message", {}).get("content", json.dumps(body)) | return body.get("answer", json.dumps(body)) | return json.dumps(body) | elif task == "image-text-to-text": - | if isinstance(body, dict) and "choices" in body: - | return body["choices"][0]["message"]["content"] + | if isinstance(body, dict) and body.get("choices"): + | return body["choices"][0].get("message", {}).get("content", json.dumps(body)) | if isinstance(body, list) and body and isinstance(body[0], dict): | return body[0].get("generated_text", json.dumps(body)) | return json.dumps(body) @@ -144,6 +148,10 @@ object ImageTaskCodegen extends TaskCodegen { | if "url" in data[0]: | return self._url_to_data_url(data[0]["url"]) | return json.dumps(body) - | elif task in ("image-classification", "object-detection", "image-segmentation", "zero-shot-image-classification"): + | elif task == "zero-shot-image-classification": + | if isinstance(body, dict) and body.get("choices"): + | return body["choices"][0].get("message", {}).get("content", json.dumps(body)) + | return json.dumps(body) + | elif task in ("image-classification", "object-detection", "image-segmentation"): | return json.dumps(body)""".stripMargin } diff --git a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala index f1806d3b94..0e026fdf0f 100644 --- a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala +++ b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala @@ -116,6 +116,53 @@ class ImageTaskCodegenSpec extends AnyFlatSpec with Matchers { out should include("json.dumps(body)") } + it should "read chat-provider responses for visual and document question-answering" in { + // #7906 Part A: these tasks only understood hf-inference's native + // {"answer": ...} shape, so a chat-completions reply was written to the + // result column as a raw JSON envelope. The native shape stays primary. + val out = ImageTaskCodegen.parsePython(makeCtx()) + val branch = out + .split("""elif task """) + .find(_.startsWith("""in ("visual-question-answering""")) + .getOrElse(fail("visual/document question-answering branch is missing")) + branch should include( + """body["choices"][0].get("message", {}).get("content", json.dumps(body))""" + ) + branch should include("""body.get("answer"""") + branch.indexOf("choices") should be < branch.indexOf("""body.get("answer"""") + } + + it should "read chat-provider responses for zero-shot-image-classification" in { + // #7906 Part A: the task used to share the image-only branch, which always + // dumps the body, so a chat reply was never extracted. It now has its own + // branch, which must precede the image-only one because the generated + // if/elif chain is first-match-wins. + val out = ImageTaskCodegen.parsePython(makeCtx()) + out should include("""elif task == "zero-shot-image-classification":""") + out should include( + """elif task in ("image-classification", "object-detection", "image-segmentation"):""" + ) + out.indexOf("""elif task == "zero-shot-image-classification":""") should be < + out.indexOf("""elif task in ("image-classification",""") + } + + it should "degrade instead of raising when a chat response is malformed" in { + // Review feedback on #7920: parsePython runs per row, so indexing straight + // into choices[0]["message"]["content"] turns one malformed provider + // response into an aborted run — an empty "choices" list raises IndexError + // and a choice without "message"/"content" raises KeyError. Every chat + // extraction in this file now uses a truthiness guard plus .get chaining, + // matching how the native shapes already degrade via .get(..., json.dumps(body)). + val out = ImageTaskCodegen.parsePython(makeCtx()) + out should not include ("""["message"]["content"]""") + out.split("""body\.get\("choices"\)""").length - 1 shouldBe 4 + out + .split( + """\.get\("message", \{\}\)\.get\("content", json\.dumps\(body\)\)""" + ) + .length - 1 shouldBe 4 + } + "ImageTaskCodegen snippets" should "never inline raw CodegenContext string values" in { // The snippets are static and reference only self.* attributes; the base // class decodes user-supplied strings safely at runtime. Sentinel values
