This is an automated email from the ASF dual-hosted git repository.
github-merge-queue[bot] pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/texera.git
The following commit(s) were added to refs/heads/main by this push:
new 378a3b54ec feat(workflow-operator): read chat-provider responses for
the image question-answering tasks (#7920)
378a3b54ec is described below
commit 378a3b54ec645e57760011bca0e08c1829c86680
Author: Prateek Ganigi <[email protected]>
AuthorDate: Thu Sep 17 22:16:29 2026 +0000
feat(workflow-operator): read chat-provider responses for the image
question-answering tasks (#7920)
### What changes were proposed in this PR?
When the operator falls back from `hf-inference` to a third-party
chat-completions provider, the reply comes back as `{"choices":
[{"message": {"content": ...}}]}`. Three image tasks in
`ImageTaskCodegen.parsePython` could not read that shape, so a correct
answer was written to the result column as a raw JSON envelope:
- `visual-question-answering` and `document-question-answering` returned
`body.get("answer", json.dumps(body))`, and a chat response has no
`answer` key.
- `zero-shot-image-classification` shared the image-only branch, which
always returns`json.dumps(body)`.
Both now read `choices[0]["message"]["content"]` when the body carries
`choices`, keeping the native `hf-inference` shape as the primary path.
`zero-shot-image-classification` gets its own branch, placed ahead of
the image-only tasks because the generated `if/elif` chain is
first-match-wins. This is the same idiom `image-to-text` and
`image-text-to-text` already use in this file, and the one applied to
the text tasks in #7798.
`image-classification`, `object-detection` and `image-segmentation` are
left as they are: they have no question to answer, so a free-text chat
reply is not meaningful structured output for them.
This is Part A of #7906 and covers the response side only. The request
side which is carrying `candidate_labels` into the chat message for
`zero-shot-image-classification`, follows in Part B.
### Any related issues?
Addresses #7906
### How was this PR tested?
133 tests pass in the `WorkflowOperator` Hugging Face suites,
`PythonCodeRawInvalidTextSpec` py-compiles the generated Python for all
117 operators, and `scalafmtCheck` is clean for main and test sources.
Two tests were added to `ImageTaskCodegenSpec`: one asserts the
visual/document question-answering branch reads `choices` ahead of the
native `answer` lookup, the other asserts the new
`zero-shot-image-classification` branch exists and precedes the
image-only branch.
The emitted Python was also exercised directly: the three fixed tasks
return the chat content, native `hf-inference` responses parse exactly
as before, non-dict and `answer`-less bodies still fall through to
`json.dumps`, and the untouched branches (`image-classification`,
`object-detection`, `image-segmentation`, `image-to-text`,
`image-text-to-text`) are unchanged.
### Was this PR authored or co-authored using generative AI tooling?
Yes, this PR was co-authored with Claude in compliance with ASF policy.
---------
Co-authored-by: Xuan Gu <[email protected]>
---
.../huggingFace/codegen/ImageTaskCodegen.scala | 18 ++++++---
.../huggingFace/codegen/ImageTaskCodegenSpec.scala | 47 ++++++++++++++++++++++
2 files changed, 60 insertions(+), 5 deletions(-)
diff --git
a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
index 673227ab9e..201d7b1623 100644
---
a/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
+++
b/common/workflow-operator/src/main/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegen.scala
@@ -106,18 +106,22 @@ object ImageTaskCodegen extends TaskCodegen {
| if isinstance(body, dict):
| if "md_results" in body:
| return body["md_results"]
- | if "choices" in body:
- | return body["choices"][0]["message"]["content"]
+ | if body.get("choices"):
+ | return body["choices"][0].get("message",
{}).get("content", json.dumps(body))
| if isinstance(body, list) and body and
isinstance(body[0], dict):
| return body[0].get("generated_text",
json.dumps(body))
| return json.dumps(body)
| elif task in ("visual-question-answering",
"document-question-answering"):
| if isinstance(body, dict):
+ | # Third-party chat providers answer via
choices[0].message;
+ | # hf-inference returns the native {"answer": ...}
shape.
+ | if body.get("choices"):
+ | return body["choices"][0].get("message",
{}).get("content", json.dumps(body))
| return body.get("answer", json.dumps(body))
| return json.dumps(body)
| elif task == "image-text-to-text":
- | if isinstance(body, dict) and "choices" in body:
- | return body["choices"][0]["message"]["content"]
+ | if isinstance(body, dict) and body.get("choices"):
+ | return body["choices"][0].get("message",
{}).get("content", json.dumps(body))
| if isinstance(body, list) and body and
isinstance(body[0], dict):
| return body[0].get("generated_text",
json.dumps(body))
| return json.dumps(body)
@@ -144,6 +148,10 @@ object ImageTaskCodegen extends TaskCodegen {
| if "url" in data[0]:
| return
self._url_to_data_url(data[0]["url"])
| return json.dumps(body)
- | elif task in ("image-classification", "object-detection",
"image-segmentation", "zero-shot-image-classification"):
+ | elif task == "zero-shot-image-classification":
+ | if isinstance(body, dict) and body.get("choices"):
+ | return body["choices"][0].get("message",
{}).get("content", json.dumps(body))
+ | return json.dumps(body)
+ | elif task in ("image-classification", "object-detection",
"image-segmentation"):
| return json.dumps(body)""".stripMargin
}
diff --git
a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
index f1806d3b94..0e026fdf0f 100644
---
a/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
+++
b/common/workflow-operator/src/test/scala/org/apache/texera/amber/operator/huggingFace/codegen/ImageTaskCodegenSpec.scala
@@ -116,6 +116,53 @@ class ImageTaskCodegenSpec extends AnyFlatSpec with
Matchers {
out should include("json.dumps(body)")
}
+ it should "read chat-provider responses for visual and document
question-answering" in {
+ // #7906 Part A: these tasks only understood hf-inference's native
+ // {"answer": ...} shape, so a chat-completions reply was written to the
+ // result column as a raw JSON envelope. The native shape stays primary.
+ val out = ImageTaskCodegen.parsePython(makeCtx())
+ val branch = out
+ .split("""elif task """)
+ .find(_.startsWith("""in ("visual-question-answering"""))
+ .getOrElse(fail("visual/document question-answering branch is missing"))
+ branch should include(
+ """body["choices"][0].get("message", {}).get("content",
json.dumps(body))"""
+ )
+ branch should include("""body.get("answer"""")
+ branch.indexOf("choices") should be <
branch.indexOf("""body.get("answer"""")
+ }
+
+ it should "read chat-provider responses for zero-shot-image-classification"
in {
+ // #7906 Part A: the task used to share the image-only branch, which always
+ // dumps the body, so a chat reply was never extracted. It now has its own
+ // branch, which must precede the image-only one because the generated
+ // if/elif chain is first-match-wins.
+ val out = ImageTaskCodegen.parsePython(makeCtx())
+ out should include("""elif task == "zero-shot-image-classification":""")
+ out should include(
+ """elif task in ("image-classification", "object-detection",
"image-segmentation"):"""
+ )
+ out.indexOf("""elif task == "zero-shot-image-classification":""") should
be <
+ out.indexOf("""elif task in ("image-classification",""")
+ }
+
+ it should "degrade instead of raising when a chat response is malformed" in {
+ // Review feedback on #7920: parsePython runs per row, so indexing straight
+ // into choices[0]["message"]["content"] turns one malformed provider
+ // response into an aborted run — an empty "choices" list raises IndexError
+ // and a choice without "message"/"content" raises KeyError. Every chat
+ // extraction in this file now uses a truthiness guard plus .get chaining,
+ // matching how the native shapes already degrade via .get(...,
json.dumps(body)).
+ val out = ImageTaskCodegen.parsePython(makeCtx())
+ out should not include ("""["message"]["content"]""")
+ out.split("""body\.get\("choices"\)""").length - 1 shouldBe 4
+ out
+ .split(
+ """\.get\("message", \{\}\)\.get\("content", json\.dumps\(body\)\)"""
+ )
+ .length - 1 shouldBe 4
+ }
+
"ImageTaskCodegen snippets" should "never inline raw CodegenContext string
values" in {
// The snippets are static and reference only self.* attributes; the base
// class decodes user-supplied strings safely at runtime. Sentinel values