The GitHub Actions job "Required Checks" on texera.git/main has failed.
Run started by GitHub user github-merge-queue[bot] (triggered by 
github-merge-queue[bot]).

Head commit for run:
ab1a8baeb67fd31b44207cf5d0575a529e1e0947 / Kary Zheng 
<[email protected]>
feat(operator): reject Count Vectorizer on the Gaussian Naive Bayes operators 
(#7599)

### What changes were proposed in this PR?

`Gaussian Naive Bayes` and `Training: Gaussian Naive Bayes` offered a
`Count Vectorizer` switch with no working setting. Turning it on ended
the execution from inside scikit-learn: `CountVectorizer` emits a sparse
matrix, and `GaussianNB` validates its input without `accept_sparse`
because it estimates a mean and a variance per feature, which reads the
zeros too. The refusal is deliberate on scikit-learn's side, since
densifying a text matrix with a large vocabulary is what would exhaust
memory.

It is specific to this estimator, not to the switch. The switch is
declared once on `SklearnModelOpDesc` and both families splice it into
the pipeline the same way, so every estimator in both groups inherits it
whether or not its own can use it.

An estimator that cannot take the sparse matrix now names what to reach
for instead, and the shared `getOutputSchemas` turns that into a
compile-time failure naming both. Nothing fires while the switch is off,
which is its default, so a freshly dropped operator is not reported
invalid before it is configured.

Hiding the switch would have been the closer match to how the family
already handles an impossible pairing, since `Tfidf Transformer` and
`Text Attribute` are hidden when `Count Vectorizer` is off. It is not
available here: a hide keys on a sibling field's value and is evaluated
in the frontend, so it cannot be made per estimator on a field the base
declares once. It would also leave a workflow already saved with the
switch on still failing at run time, which the compile-time check
catches.

### Any related issues, documentation, discussions?

Fixes #7597.

### How was this PR tested?

- The premise checked against scikit-learn: fitting
`make_pipeline(CountVectorizer(), Estimator())` on the same documents,
`GaussianNB` raises the TypeError in the issue while `MultinomialNB`,
`BernoulliNB` and `ComplementNB` all fit, which is what the message
points at.
- Both operators' specs assert the rejection carries the estimator's own
name and the alternatives.
- The base's spec covers the three states that matter: an estimator
naming no alternative still accepts the switch, one naming an
alternative rejects it, and neither fires while the switch is off.
- `WorkflowOperator/testOnly *Sklearn*`: 62 suites, 421 tests, none
failed. `scalafmtCheckAll` and both scalafix checks clean.

### Was this PR authored or co-authored using generative AI tooling?

Generated-by: Claude Code (Claude Opus 5)

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>

Report URL: https://github.com/apache/texera/actions/runs/32798712137

With regards,
GitHub Actions via GitBox

Reply via email to