[ 
https://issues.apache.org/jira/browse/CAMEL-25311?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Luigi De Masi updated CAMEL-25311:
----------------------------------
    Description: 
h2. Problem and intended behaviour

Add a dedicated Wolf-Defender semantic expert in 
{{components/camel-ai/camel-wolf-defender}}, published as 
{{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen untrusted 
text for prompt injection through {{camel-semantic}}.

For example, a route receiving a retrieved document or tool response should be 
able to evaluate {{ref:injection}} and route the exchange to processing, 
rejection, or review according to application policy. The route should not need 
to know the model's label IDs, tokenizer, tensor format, inference library, or 
document-scoring implementation.

This is the concrete provider implementation associated with CAMEL-25310, which 
introduces expert selection, capability reporting, optional instructions, and 
catalog metadata. Reuse those contracts rather than implementing a separate 
provider-selection mechanism in this module. Use *expert* in configuration and 
documentation; implementing the existing {{SemanticAdapter}} SPI does not 
require renaming that Java API.

h2. Scope and packaging

* Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate 
its dependencies, service discovery, generated metadata, documentation, and 
build registration using Camel conventions.
* Implement a dedicated semantic adapter/expert for the fixed prompt-injection 
evaluation.
* Keep {{camel-semantic}} independent of concrete inference and tokenization 
libraries. These dependencies belong to the optional Wolf-Defender module.
* Do not depend on LangChain4j or require a Jev-compatible service. 
Applications should keep the common semantic contract when combining this 
expert with other providers.
* Target local inference for the initial implementation. ONNX Runtime's Java 
API is a candidate backend; select and document the runtime/tokenizer 
combination after checking the chosen export. A generic backend framework, 
remote model-serving protocol, and separate routing endpoint are not 
prerequisites.
* Establish a tested baseline with a pinned Wolf-Defender v2 artifact, such as 
Small. Document the supported model/export/runtime combinations; do not imply 
that all variants and quantizations have been validated.

h2. Input and capability contract

Expose the following capabilities through the framework from CAMEL-25310:

||Property||Required contract||
|Input|Text selected by the semantic declaration's state expression|
|Result type|BOOLEAN|
|Positive meaning|Prompt injection or jailbreak-like instruction detected|
|Probability|Probability assigned to the INJECTION class|
|Instructions|Unsupported; no question or synthetic instruction prompt is 
required|
|Caller-defined criteria|Unsupported|
|CHOICE / SCORE|Unsupported|

The dedicated expert supplies the meaning of the evaluation. No new {{task}} 
keyword is needed, and users do not need to repeat a question such as "Is this 
a prompt injection?".

Reject CHOICE/SCORE declarations, unsupported instructions, and arbitrary 
criteria during declaration validation, before inference. Errors should 
identify the named evaluation and selected expert. Never ignore unsupported 
options or silently select a different provider. Validate the selected runtime 
value as text; structured objects require explicit application selection or 
conversion rather than an implicit {{toString()}}.

Preserve the exchange content and existing semantic result publication 
behaviour. The expert's result should not replace the original message body 
with a model-specific response.

h2. Mapping model output to SemanticResult

The current Wolf-Defender Small model card defines class 0 as BENIGN and class 
1 as INJECTION. Its ONNX example returns logits. The adapter must validate the 
selected artifact's output contract and map the positive class correctly.

* Convert two-class logits into probabilities using numerically stable softmax, 
where that is the selected export's contract. Do not apply softmax a second 
time to an already normalized probability output.
* Return {{P(INJECTION)}} through the existing BOOLEAN probability 
representation. For example, a top-label response equivalent to BENIGN with 
probability 0.97 must yield an injection probability of 0.03, not 0.97.
* Validate output dimensions, label mapping, finite numeric values, and 
probability bounds. Incompatible artifacts and malformed outputs are errors.
* Preserve probability information until the common semantic threshold and 
uncertainty policy are applied. Do not collapse it to the model's default label 
first or introduce a conflicting hidden decision threshold.
* Do not advertise this probability as an application-defined SCORE severity 
scale, or invent a calibrated confidence guarantee.

BENIGN means that this evaluation did not detect injection. It does not 
establish general safety, authorization, or the absence of other threats. 
Actions such as block, quarantine, review, or continue belong to Camel routing 
policy; they are not additional model labels.

h2. Tokenization and document handling

Own tokenization, input tensors, special tokens, attention masks, padding, and 
model invocation inside the expert. Use tokenizer assets compatible with the 
pinned model and validate against reference inference.

The Small v2 model card describes a 2,048-token window. Its published document 
evaluation uses overlapping windows with 64-token overlap and normalized 
Smooth-Max aggregation. A single truncated window does not reproduce that 
protocol.

Define and document bounded handling of longer input. Either implement a 
documented document-scoring policy or reject inputs beyond the supported limit 
explicitly; never silently discard an unchecked suffix. If adopting the 
published aggregation, verify the exact formula and parameters against an 
authoritative implementation rather than guessing from its name. Clearly 
identify any alternative policy and its implications for threshold selection.

For supported multi-window evaluation, test injection-bearing content beyond 
the first window and at window boundaries. Limits should bound both 
tokenization work and the number of inference windows. Empty, missing, and 
oversized input must have an explicit, tested outcome.

h2. Expert configuration and lifecycle

Keep model-specific configuration on the configured Wolf expert instance: 
model/tokenizer locations and revision, selected export, runtime options, and 
input/document limits. Named evaluations retain common semantic options such as 
state, threshold, and uncertainty policy.

Support reproducible use of explicitly provisioned local artifacts without 
requiring downloads during normal route execution. Model weights should not be 
bundled into the Camel source repository or ordinary module artifact. Document 
acquisition, artifact identity, and compatible tokenizer/configuration files.

Load and reuse model resources through Camel-managed lifecycle. Separate 
declaration/capability inspection from model loading and inference. Close 
sessions, tensors, and tokenizer/native resources on shutdown and startup 
failure. Support concurrent evaluations without cross-exchange state leakage, 
and bound concurrency and resource consumption using the selected runtime's 
facilities.

Respect the semantic SPI's timeout, interruption, and shutdown contract. Verify 
the actual native-runtime cancellation behaviour; a timeout around a Java task 
must not be described as stopping native inference if it only stops waiting for 
it.

Model-loading failures, invalid inputs, inference failures, malformed output, 
and cancellation must remain errors. They must not become a benign result or an 
injection probability of zero. Diagnostics should include enough 
provider/artifact identity to reproduce a problem without logging submitted 
text or credentials.

h2. Integration with experts and Camel Catalog

Register the expert through the discovery/lifecycle mechanism agreed in 
CAMEL-25310 and allow explicitly configured registry instances. Reuse its 
explicit/default/sole-expert resolution rules and its error on ambiguity. Two 
configured Wolf instances may use different artifacts or limits and must retain 
their separate identities.

Publish generated static capability metadata for {{wolf-defender}}, linked to 
{{camel-wolf-defender}}, through the same authoritative definition used by 
runtime capabilities. Catalog inspection must work without downloading weights 
or initializing a native runtime. Static metadata describes the provider's 
contract; it must not claim that an arbitrary local artifact has been validated.

Support the existing single-evaluation and batch SPI contracts. Framework 
grouping across experts remains the responsibility of CAMEL-25310. The initial 
implementation may use the existing sequential batch behaviour; optimized 
tensor batching is not required. Preserve named result keys and all-or-error 
publication.

h2. Illustrative application declaration

The following uses the syntax proposed in CAMEL-25310 and assumes a configured 
Wolf expert instance named {{security}}. It is illustrative, pending the final 
framework API and binding conventions.

{noformat}
- semantic:
    question:
      injection:
        expert: security
        type: boolean
        state: "${body}"
        threshold: "{{security.injection.threshold}}"
        uncertainty: "{{security.injection.uncertainty}}"
        uncertaintyPolicy: fail
{noformat}

Routes use {{ref:injection}} through the semantic language. Provide a runnable 
example showing expert construction/configuration, artifact provisioning, and 
routing on the resulting decision. Demonstrate that model-specific options 
remain outside the named evaluation. Explain how uncertainty and operational 
errors are handled by the route.

h2. License and documentation

The Small model card declares Apache License 2.0 and retains MIT terms for the 
upstream mmBERT-small portions. Preserve applicable notices for any 
redistributed material and document model provenance separately from 
runtime/tokenizer dependency licenses. Confirm the license of the exact 
selected artifacts as part of implementation.

Document installation, supported runtime/platform combinations, configuration, 
fixed BOOLEAN semantics, probability mapping, limits, and error behaviour. 
Describe false-positive/false-negative limitations without making benchmark or 
latency claims for an untested Java implementation.

h2. Validation and acceptance criteria

# The new {{components/camel-ai/camel-wolf-defender}} module builds and 
integrates with Camel's normal generated metadata and documentation processes.
# A named BOOLEAN evaluation without instructions works through 
{{camel-semantic}}, with both explicit expert selection and the framework's 
sole-expert discovery behaviour.
# Unsupported result types, instructions, and criteria fail during declaration 
validation without running inference. Unknown/ambiguous expert selection 
follows CAMEL-25310.
# Tests verify positive-class mapping, stable probability conversion, 
malformed/non-finite output rejection, and common threshold/uncertainty 
boundaries, including a high-confidence BENIGN result.
# Input validation and document limits are explicit and tested. Oversized input 
is never silently partially screened; any supported windowed policy has 
boundary and late-window coverage.
# Lifecycle and failure tests cover resource cleanup, concurrent use, 
cancellation/timeout behaviour, and errors remaining distinct from negative 
classifications.
# Batch evaluation preserves keys, expert-instance identity, and complete 
result validation without exposing partial success.
# Static capabilities are discoverable through Camel Catalog, agree with 
runtime reporting, and can be read without model/native-runtime initialization.
# Normal unit tests use deterministic fixtures and require neither network 
access nor a large model download. Add an opt-in integration test using a 
pinned real artifact to verify tokenizer/inference parity against reference 
outputs with a documented numeric tolerance.
# Provide a runnable local example and setup instructions, including the tested 
artifact/runtime, license/provenance, and representative benign, injection, and 
difficult-benign inputs. Keep adapter correctness evidence separate from claims 
about model detection quality.

h2. Related issue and references

* [CAMEL-25310: semantic experts, capabilities and catalog 
metadata|https://issues.apache.org/jira/browse/CAMEL-25310]
* [Wolf-Defender Small model card and inference 
examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
* [Wolf-Defender full 
model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
* [Wolf-Defender Small license 
file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
* [Wolf-Defender v2 
article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
* [ONNX Runtime Java API|https://onnxruntime.ai/docs/get-started/with-java.html]


  was:
h2. Problem and intended behaviour

Add a dedicated Wolf-Defender semantic expert in 
{{components/camel-ai/camel-wolf-defender}}, published as 
{{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen untrusted 
text for prompt injection through {{camel-semantic}}.

For example, a route receiving a retrieved document or tool response should be 
able to evaluate {{ref:injection}} and route the exchange to processing, 
rejection, or review according to application policy. The route should not need 
to know the model's label IDs, tokenizer, tensor format, inference library, or 
document-scoring implementation.

This is the concrete provider implementation associated with CAMEL-25310, which 
introduces expert selection, capability reporting, optional instructions, and 
catalog metadata. Reuse those contracts rather than implementing a separate 
provider-selection mechanism in this module. Use *expert* in configuration and 
documentation; implementing the existing {{SemanticAdapter}} SPI does not 
require renaming that Java API.

h2. Scope and packaging

* Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate 
its dependencies, service discovery, generated metadata, documentation, and 
build registration using Camel conventions.
* Implement a dedicated semantic adapter/expert for the fixed prompt-injection 
evaluation.
* Keep {{camel-semantic}} independent of concrete inference and tokenization 
libraries. These dependencies belong to the optional Wolf-Defender module.
* Do not depend on LangChain4j or require a Jev-compatible service. 
Applications should keep the common semantic contract when combining this 
expert with other providers.
* Target local inference for the initial implementation. ONNX Runtime's Java 
API is a candidate backend; select and document the runtime/tokenizer 
combination after checking the chosen export. A generic backend framework, 
remote model-serving protocol, and separate routing endpoint are not 
prerequisites.
* Establish a tested baseline with a pinned Wolf-Defender v2 artifact, such as 
Small. Document the supported model/export/runtime combinations; do not imply 
that all variants and quantizations have been validated.

h2. Input and capability contract

Expose the following capabilities through the framework from CAMEL-25310:

||Property||Required contract||
|Input|Text selected by the semantic declaration's state expression|
|Result type|BOOLEAN|
|Positive meaning|Prompt injection or jailbreak-like instruction detected|
|Probability|Probability assigned to the INJECTION class|
|Instructions|Unsupported; no question or synthetic instruction prompt is 
required|
|Caller-defined criteria|Unsupported|
|CHOICE / SCORE|Unsupported|

The dedicated expert supplies the meaning of the evaluation. No new {{task}} 
keyword is needed, and users do not need to repeat a question such as "Is this 
a prompt injection?".

Reject CHOICE/SCORE declarations, unsupported instructions, and arbitrary 
criteria during declaration validation, before inference. Errors should 
identify the named evaluation and selected expert. Never ignore unsupported 
options or silently select a different provider. Validate the selected runtime 
value as text; structured objects require explicit application selection or 
conversion rather than an implicit {{toString()}}.

Preserve the exchange content and existing semantic result publication 
behaviour. The expert's result should not replace the original message body 
with a model-specific response.

h2. Mapping model output to SemanticResult

The current Wolf-Defender Small model card defines class 0 as BENIGN and class 
1 as INJECTION. Its ONNX example returns logits. The adapter must validate the 
selected artifact's output contract and map the positive class correctly.

* Convert two-class logits into probabilities using numerically stable softmax, 
where that is the selected export's contract. Do not apply softmax a second 
time to an already normalized probability output.
* Return {{P(INJECTION)}} through the existing BOOLEAN probability 
representation. For example, a top-label response equivalent to BENIGN with 
probability 0.97 must yield an injection probability of 0.03, not 0.97.
* Validate output dimensions, label mapping, finite numeric values, and 
probability bounds. Incompatible artifacts and malformed outputs are errors.
* Preserve probability information until the common semantic threshold and 
uncertainty policy are applied. Do not collapse it to the model's default label 
first or introduce a conflicting hidden decision threshold.
* Do not advertise this probability as an application-defined SCORE severity 
scale, or invent a calibrated confidence guarantee.

BENIGN means that this evaluation did not detect injection. It does not 
establish general safety, authorization, or the absence of other threats. 
Actions such as block, quarantine, review, or continue belong to Camel routing 
policy; they are not additional model labels.

h2. Tokenization and document handling

Own tokenization, input tensors, special tokens, attention masks, padding, and 
model invocation inside the expert. Use tokenizer assets compatible with the 
pinned model and validate against reference inference.

The Small v2 model card describes a 2,048-token window. Its published document 
evaluation uses overlapping windows with 64-token overlap and normalized 
Smooth-Max aggregation. A single truncated window does not reproduce that 
protocol.

Define and document bounded handling of longer input. Either implement a 
documented document-scoring policy or reject inputs beyond the supported limit 
explicitly; never silently discard an unchecked suffix. If adopting the 
published aggregation, verify the exact formula and parameters against an 
authoritative implementation rather than guessing from its name. Clearly 
identify any alternative policy and its implications for threshold selection.

For supported multi-window evaluation, test injection-bearing content beyond 
the first window and at window boundaries. Limits should bound both 
tokenization work and the number of inference windows. Empty, missing, and 
oversized input must have an explicit, tested outcome.

h2. Expert configuration and lifecycle

Keep model-specific configuration on the configured Wolf expert instance: 
model/tokenizer locations and revision, selected export, runtime options, and 
input/document limits. Named evaluations retain common semantic options such as 
state, threshold, and uncertainty policy.

Support reproducible use of explicitly provisioned local artifacts without 
requiring downloads during normal route execution. Model weights should not be 
bundled into the Camel source repository or ordinary module artifact. Document 
acquisition, artifact identity, and compatible tokenizer/configuration files.

Load and reuse model resources through Camel-managed lifecycle. Separate 
declaration/capability inspection from model loading and inference. Close 
sessions, tensors, and tokenizer/native resources on shutdown and startup 
failure. Support concurrent evaluations without cross-exchange state leakage, 
and bound concurrency and resource consumption using the selected runtime's 
facilities.

Respect the semantic SPI's timeout, interruption, and shutdown contract. Verify 
the actual native-runtime cancellation behaviour; a timeout around a Java task 
must not be described as stopping native inference if it only stops waiting for 
it.

Model-loading failures, invalid inputs, inference failures, malformed output, 
and cancellation must remain errors. They must not become a benign result or an 
injection probability of zero. Diagnostics should include enough 
provider/artifact identity to reproduce a problem without logging submitted 
text or credentials.

h2. Integration with experts and Camel Catalog

Register the expert through the discovery/lifecycle mechanism agreed in 
CAMEL-25310 and allow explicitly configured registry instances. Reuse its 
explicit/default/sole-expert resolution rules and its error on ambiguity. Two 
configured Wolf instances may use different artifacts or limits and must retain 
their separate identities.

Publish generated static capability metadata for {{wolf-defender}}, linked to 
{{camel-wolf-defender}}, through the same authoritative definition used by 
runtime capabilities. Catalog inspection must work without downloading weights 
or initializing a native runtime. Static metadata describes the provider's 
contract; it must not claim that an arbitrary local artifact has been validated.

Support the existing single-evaluation and batch SPI contracts. Framework 
grouping across experts remains the responsibility of CAMEL-25310. The initial 
implementation may use the existing sequential batch behaviour; optimized 
tensor batching is not required. Preserve named result keys and all-or-error 
publication.

h2. Illustrative application declaration

The following uses the syntax proposed in CAMEL-25310 and assumes a configured 
Wolf expert instance named {{security}}. It is illustrative, pending the final 
framework API and binding conventions.

{code:yaml}
- semantic:
    question:
      injection:
        expert: security
        type: boolean
        state: "${body}"
        threshold: "{{security.injection.threshold}}"
        uncertainty: "{{security.injection.uncertainty}}"
        uncertaintyPolicy: fail
{code}

Routes use {{ref:injection}} through the semantic language. Provide a runnable 
example showing expert construction/configuration, artifact provisioning, and 
routing on the resulting decision. Demonstrate that model-specific options 
remain outside the named evaluation. Explain how uncertainty and operational 
errors are handled by the route.

h2. License and documentation

The Small model card declares Apache License 2.0 and retains MIT terms for the 
upstream mmBERT-small portions. Preserve applicable notices for any 
redistributed material and document model provenance separately from 
runtime/tokenizer dependency licenses. Confirm the license of the exact 
selected artifacts as part of implementation.

Document installation, supported runtime/platform combinations, configuration, 
fixed BOOLEAN semantics, probability mapping, limits, and error behaviour. 
Describe false-positive/false-negative limitations without making benchmark or 
latency claims for an untested Java implementation.

h2. Validation and acceptance criteria

# The new {{components/camel-ai/camel-wolf-defender}} module builds and 
integrates with Camel's normal generated metadata and documentation processes.
# A named BOOLEAN evaluation without instructions works through 
{{camel-semantic}}, with both explicit expert selection and the framework's 
sole-expert discovery behaviour.
# Unsupported result types, instructions, and criteria fail during declaration 
validation without running inference. Unknown/ambiguous expert selection 
follows CAMEL-25310.
# Tests verify positive-class mapping, stable probability conversion, 
malformed/non-finite output rejection, and common threshold/uncertainty 
boundaries, including a high-confidence BENIGN result.
# Input validation and document limits are explicit and tested. Oversized input 
is never silently partially screened; any supported windowed policy has 
boundary and late-window coverage.
# Lifecycle and failure tests cover resource cleanup, concurrent use, 
cancellation/timeout behaviour, and errors remaining distinct from negative 
classifications.
# Batch evaluation preserves keys, expert-instance identity, and complete 
result validation without exposing partial success.
# Static capabilities are discoverable through Camel Catalog, agree with 
runtime reporting, and can be read without model/native-runtime initialization.
# Normal unit tests use deterministic fixtures and require neither network 
access nor a large model download. Add an opt-in integration test using a 
pinned real artifact to verify tokenizer/inference parity against reference 
outputs with a documented numeric tolerance.
# Provide a runnable local example and setup instructions, including the tested 
artifact/runtime, license/provenance, and representative benign, injection, and 
difficult-benign inputs. Keep adapter correctness evidence separate from claims 
about model detection quality.

h2. Related issue and references

* [CAMEL-25310: semantic experts, capabilities and catalog 
metadata|https://issues.apache.org/jira/browse/CAMEL-25310]
* [Wolf-Defender Small model card and inference 
examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
* [Wolf-Defender full 
model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
* [Wolf-Defender Small license 
file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
* [Wolf-Defender v2 
article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
* [ONNX Runtime Java API|https://onnxruntime.ai/docs/get-started/with-java.html]



> camel-wolf-defender: Add a semantic expert for prompt-injection detection
> -------------------------------------------------------------------------
>
>                 Key: CAMEL-25311
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25311
>             Project: Camel
>          Issue Type: New Feature
>          Components: camel-ai
>            Reporter: Luigi De Masi
>            Assignee: Luigi De Masi
>            Priority: Major
>
> h2. Problem and intended behaviour
> Add a dedicated Wolf-Defender semantic expert in 
> {{components/camel-ai/camel-wolf-defender}}, published as 
> {{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen 
> untrusted text for prompt injection through {{camel-semantic}}.
> For example, a route receiving a retrieved document or tool response should 
> be able to evaluate {{ref:injection}} and route the exchange to processing, 
> rejection, or review according to application policy. The route should not 
> need to know the model's label IDs, tokenizer, tensor format, inference 
> library, or document-scoring implementation.
> This is the concrete provider implementation associated with CAMEL-25310, 
> which introduces expert selection, capability reporting, optional 
> instructions, and catalog metadata. Reuse those contracts rather than 
> implementing a separate provider-selection mechanism in this module. Use 
> *expert* in configuration and documentation; implementing the existing 
> {{SemanticAdapter}} SPI does not require renaming that Java API.
> h2. Scope and packaging
> * Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate 
> its dependencies, service discovery, generated metadata, documentation, and 
> build registration using Camel conventions.
> * Implement a dedicated semantic adapter/expert for the fixed 
> prompt-injection evaluation.
> * Keep {{camel-semantic}} independent of concrete inference and tokenization 
> libraries. These dependencies belong to the optional Wolf-Defender module.
> * Do not depend on LangChain4j or require a Jev-compatible service. 
> Applications should keep the common semantic contract when combining this 
> expert with other providers.
> * Target local inference for the initial implementation. ONNX Runtime's Java 
> API is a candidate backend; select and document the runtime/tokenizer 
> combination after checking the chosen export. A generic backend framework, 
> remote model-serving protocol, and separate routing endpoint are not 
> prerequisites.
> * Establish a tested baseline with a pinned Wolf-Defender v2 artifact, such 
> as Small. Document the supported model/export/runtime combinations; do not 
> imply that all variants and quantizations have been validated.
> h2. Input and capability contract
> Expose the following capabilities through the framework from CAMEL-25310:
> ||Property||Required contract||
> |Input|Text selected by the semantic declaration's state expression|
> |Result type|BOOLEAN|
> |Positive meaning|Prompt injection or jailbreak-like instruction detected|
> |Probability|Probability assigned to the INJECTION class|
> |Instructions|Unsupported; no question or synthetic instruction prompt is 
> required|
> |Caller-defined criteria|Unsupported|
> |CHOICE / SCORE|Unsupported|
> The dedicated expert supplies the meaning of the evaluation. No new {{task}} 
> keyword is needed, and users do not need to repeat a question such as "Is 
> this a prompt injection?".
> Reject CHOICE/SCORE declarations, unsupported instructions, and arbitrary 
> criteria during declaration validation, before inference. Errors should 
> identify the named evaluation and selected expert. Never ignore unsupported 
> options or silently select a different provider. Validate the selected 
> runtime value as text; structured objects require explicit application 
> selection or conversion rather than an implicit {{toString()}}.
> Preserve the exchange content and existing semantic result publication 
> behaviour. The expert's result should not replace the original message body 
> with a model-specific response.
> h2. Mapping model output to SemanticResult
> The current Wolf-Defender Small model card defines class 0 as BENIGN and 
> class 1 as INJECTION. Its ONNX example returns logits. The adapter must 
> validate the selected artifact's output contract and map the positive class 
> correctly.
> * Convert two-class logits into probabilities using numerically stable 
> softmax, where that is the selected export's contract. Do not apply softmax a 
> second time to an already normalized probability output.
> * Return {{P(INJECTION)}} through the existing BOOLEAN probability 
> representation. For example, a top-label response equivalent to BENIGN with 
> probability 0.97 must yield an injection probability of 0.03, not 0.97.
> * Validate output dimensions, label mapping, finite numeric values, and 
> probability bounds. Incompatible artifacts and malformed outputs are errors.
> * Preserve probability information until the common semantic threshold and 
> uncertainty policy are applied. Do not collapse it to the model's default 
> label first or introduce a conflicting hidden decision threshold.
> * Do not advertise this probability as an application-defined SCORE severity 
> scale, or invent a calibrated confidence guarantee.
> BENIGN means that this evaluation did not detect injection. It does not 
> establish general safety, authorization, or the absence of other threats. 
> Actions such as block, quarantine, review, or continue belong to Camel 
> routing policy; they are not additional model labels.
> h2. Tokenization and document handling
> Own tokenization, input tensors, special tokens, attention masks, padding, 
> and model invocation inside the expert. Use tokenizer assets compatible with 
> the pinned model and validate against reference inference.
> The Small v2 model card describes a 2,048-token window. Its published 
> document evaluation uses overlapping windows with 64-token overlap and 
> normalized Smooth-Max aggregation. A single truncated window does not 
> reproduce that protocol.
> Define and document bounded handling of longer input. Either implement a 
> documented document-scoring policy or reject inputs beyond the supported 
> limit explicitly; never silently discard an unchecked suffix. If adopting the 
> published aggregation, verify the exact formula and parameters against an 
> authoritative implementation rather than guessing from its name. Clearly 
> identify any alternative policy and its implications for threshold selection.
> For supported multi-window evaluation, test injection-bearing content beyond 
> the first window and at window boundaries. Limits should bound both 
> tokenization work and the number of inference windows. Empty, missing, and 
> oversized input must have an explicit, tested outcome.
> h2. Expert configuration and lifecycle
> Keep model-specific configuration on the configured Wolf expert instance: 
> model/tokenizer locations and revision, selected export, runtime options, and 
> input/document limits. Named evaluations retain common semantic options such 
> as state, threshold, and uncertainty policy.
> Support reproducible use of explicitly provisioned local artifacts without 
> requiring downloads during normal route execution. Model weights should not 
> be bundled into the Camel source repository or ordinary module artifact. 
> Document acquisition, artifact identity, and compatible 
> tokenizer/configuration files.
> Load and reuse model resources through Camel-managed lifecycle. Separate 
> declaration/capability inspection from model loading and inference. Close 
> sessions, tensors, and tokenizer/native resources on shutdown and startup 
> failure. Support concurrent evaluations without cross-exchange state leakage, 
> and bound concurrency and resource consumption using the selected runtime's 
> facilities.
> Respect the semantic SPI's timeout, interruption, and shutdown contract. 
> Verify the actual native-runtime cancellation behaviour; a timeout around a 
> Java task must not be described as stopping native inference if it only stops 
> waiting for it.
> Model-loading failures, invalid inputs, inference failures, malformed output, 
> and cancellation must remain errors. They must not become a benign result or 
> an injection probability of zero. Diagnostics should include enough 
> provider/artifact identity to reproduce a problem without logging submitted 
> text or credentials.
> h2. Integration with experts and Camel Catalog
> Register the expert through the discovery/lifecycle mechanism agreed in 
> CAMEL-25310 and allow explicitly configured registry instances. Reuse its 
> explicit/default/sole-expert resolution rules and its error on ambiguity. Two 
> configured Wolf instances may use different artifacts or limits and must 
> retain their separate identities.
> Publish generated static capability metadata for {{wolf-defender}}, linked to 
> {{camel-wolf-defender}}, through the same authoritative definition used by 
> runtime capabilities. Catalog inspection must work without downloading 
> weights or initializing a native runtime. Static metadata describes the 
> provider's contract; it must not claim that an arbitrary local artifact has 
> been validated.
> Support the existing single-evaluation and batch SPI contracts. Framework 
> grouping across experts remains the responsibility of CAMEL-25310. The 
> initial implementation may use the existing sequential batch behaviour; 
> optimized tensor batching is not required. Preserve named result keys and 
> all-or-error publication.
> h2. Illustrative application declaration
> The following uses the syntax proposed in CAMEL-25310 and assumes a 
> configured Wolf expert instance named {{security}}. It is illustrative, 
> pending the final framework API and binding conventions.
> {noformat}
> - semantic:
>     question:
>       injection:
>         expert: security
>         type: boolean
>         state: "${body}"
>         threshold: "{{security.injection.threshold}}"
>         uncertainty: "{{security.injection.uncertainty}}"
>         uncertaintyPolicy: fail
> {noformat}
> Routes use {{ref:injection}} through the semantic language. Provide a 
> runnable example showing expert construction/configuration, artifact 
> provisioning, and routing on the resulting decision. Demonstrate that 
> model-specific options remain outside the named evaluation. Explain how 
> uncertainty and operational errors are handled by the route.
> h2. License and documentation
> The Small model card declares Apache License 2.0 and retains MIT terms for 
> the upstream mmBERT-small portions. Preserve applicable notices for any 
> redistributed material and document model provenance separately from 
> runtime/tokenizer dependency licenses. Confirm the license of the exact 
> selected artifacts as part of implementation.
> Document installation, supported runtime/platform combinations, 
> configuration, fixed BOOLEAN semantics, probability mapping, limits, and 
> error behaviour. Describe false-positive/false-negative limitations without 
> making benchmark or latency claims for an untested Java implementation.
> h2. Validation and acceptance criteria
> # The new {{components/camel-ai/camel-wolf-defender}} module builds and 
> integrates with Camel's normal generated metadata and documentation processes.
> # A named BOOLEAN evaluation without instructions works through 
> {{camel-semantic}}, with both explicit expert selection and the framework's 
> sole-expert discovery behaviour.
> # Unsupported result types, instructions, and criteria fail during 
> declaration validation without running inference. Unknown/ambiguous expert 
> selection follows CAMEL-25310.
> # Tests verify positive-class mapping, stable probability conversion, 
> malformed/non-finite output rejection, and common threshold/uncertainty 
> boundaries, including a high-confidence BENIGN result.
> # Input validation and document limits are explicit and tested. Oversized 
> input is never silently partially screened; any supported windowed policy has 
> boundary and late-window coverage.
> # Lifecycle and failure tests cover resource cleanup, concurrent use, 
> cancellation/timeout behaviour, and errors remaining distinct from negative 
> classifications.
> # Batch evaluation preserves keys, expert-instance identity, and complete 
> result validation without exposing partial success.
> # Static capabilities are discoverable through Camel Catalog, agree with 
> runtime reporting, and can be read without model/native-runtime 
> initialization.
> # Normal unit tests use deterministic fixtures and require neither network 
> access nor a large model download. Add an opt-in integration test using a 
> pinned real artifact to verify tokenizer/inference parity against reference 
> outputs with a documented numeric tolerance.
> # Provide a runnable local example and setup instructions, including the 
> tested artifact/runtime, license/provenance, and representative benign, 
> injection, and difficult-benign inputs. Keep adapter correctness evidence 
> separate from claims about model detection quality.
> h2. Related issue and references
> * [CAMEL-25310: semantic experts, capabilities and catalog 
> metadata|https://issues.apache.org/jira/browse/CAMEL-25310]
> * [Wolf-Defender Small model card and inference 
> examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
> * [Wolf-Defender full 
> model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
> * [Wolf-Defender Small license 
> file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
> * [Wolf-Defender v2 
> article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
> * [ONNX Runtime Java 
> API|https://onnxruntime.ai/docs/get-started/with-java.html]



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to