[
https://issues.apache.org/jira/browse/CAMEL-25311?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123470#comment-18123470
]
Claus Ibsen commented on CAMEL-25311:
-------------------------------------
Related to observability for a guardrail-type expert like this one, and not a
blocker for the initial implementation.
OpenTelemetry is working on guardrail-specific conventions:
semantic-conventions-genai PR #427 "run guardrail span and security finding"
(still open, so names may change):
https://github.com/open-telemetry/semantic-conventions-genai/pull/427
It adds {{gen_ai.run_guardrail.internal}} (in-process checks, which fits a
local ONNX expert) and a {{gen_ai.security.finding}} event with:
* {{gen_ai.security.guardrail.*}}: which guardrail ran (the expert and pinned
artifact identity);
* {{gen_ai.security.verdict.*}}: what it returned (BENIGN/INJECTION);
* {{gen_ai.security.risk.*}}: classification and score (P(INJECTION); a
category such as OWASP LLM01 Prompt Injection);
* {{gen_ai.security.action.type}}: what the *caller* enforced, kept separate
from the verdict;
* {{gen_ai.security.policy.*}} / {{policy.rule.id}}: which policy fired.
The verdict/action split matches this issue's rule that the model returns a
verdict while block / quarantine / review / continue belongs to Camel routing
policy. It also fits the requirement not to log the submitted text: none of
these attributes needs the input. Worth keeping in mind for the capability
descriptor and result shape, so the expert can emit these once #427 lands.
_Claude Code on behalf of davsclaus_
> camel-wolf-defender: Add a semantic expert for prompt-injection detection
> -------------------------------------------------------------------------
>
> Key: CAMEL-25311
> URL: https://issues.apache.org/jira/browse/CAMEL-25311
> Project: Camel
> Issue Type: New Feature
> Components: camel-ai
> Reporter: Luigi De Masi
> Assignee: Luigi De Masi
> Priority: Major
>
> h2. Problem and intended behaviour
> Add a dedicated Wolf-Defender semantic expert in
> {{components/camel-ai/camel-wolf-defender}}, published as
> {{org.apache.camel:camel-wolf-defender}}, so Camel routes can screen
> untrusted text for prompt injection through {{camel-semantic}}.
> For example, a route receiving a retrieved document or tool response should
> be able to evaluate {{ref:injection}} and route the exchange to processing,
> rejection, or review according to application policy. The route should not
> need to know the model's label IDs, tokenizer, tensor format, inference
> library, or document-scoring implementation.
> This is the concrete provider implementation associated with CAMEL-25310,
> which introduces expert selection, capability reporting, optional
> instructions, and catalog metadata. Reuse those contracts rather than
> implementing a separate provider-selection mechanism in this module. Use
> *expert* in configuration and documentation; implementing the existing
> {{SemanticAdapter}} SPI does not require renaming that Java API.
> h2. Scope and packaging
> * Create {{camel-wolf-defender}} under the {{camel-ai}} parent and integrate
> its dependencies, service discovery, generated metadata, documentation, and
> build registration using Camel conventions.
> * Implement a dedicated semantic adapter/expert for the fixed
> prompt-injection evaluation.
> * Keep {{camel-semantic}} independent of concrete inference and tokenization
> libraries. These dependencies belong to the optional Wolf-Defender module.
> * Do not depend on LangChain4j or require a Jev-compatible service.
> Applications should keep the common semantic contract when combining this
> expert with other providers.
> * Target local inference for the initial implementation. ONNX Runtime's Java
> API is a candidate backend; select and document the runtime/tokenizer
> combination after checking the chosen export. A generic backend framework,
> remote model-serving protocol, and separate routing endpoint are not
> prerequisites.
> * Establish a tested baseline with a pinned Wolf-Defender v2 artifact, such
> as Small. Document the supported model/export/runtime combinations; do not
> imply that all variants and quantizations have been validated.
> h2. Input and capability contract
> Expose the following capabilities through the framework from CAMEL-25310:
> ||Property||Required contract||
> |Input|Text selected by the semantic declaration's state expression|
> |Result type|BOOLEAN|
> |Positive meaning|Prompt injection or jailbreak-like instruction detected|
> |Probability|Probability assigned to the INJECTION class|
> |Instructions|Unsupported; no question or synthetic instruction prompt is
> required|
> |Caller-defined criteria|Unsupported|
> |CHOICE / SCORE|Unsupported|
> The dedicated expert supplies the meaning of the evaluation. No new {{task}}
> keyword is needed, and users do not need to repeat a question such as "Is
> this a prompt injection?".
> Reject CHOICE/SCORE declarations, unsupported instructions, and arbitrary
> criteria during declaration validation, before inference. Errors should
> identify the named evaluation and selected expert. Never ignore unsupported
> options or silently select a different provider. Validate the selected
> runtime value as text; structured objects require explicit application
> selection or conversion rather than an implicit {{toString()}}.
> Preserve the exchange content and existing semantic result publication
> behaviour. The expert's result should not replace the original message body
> with a model-specific response.
> h2. Mapping model output to SemanticResult
> The current Wolf-Defender Small model card defines class 0 as BENIGN and
> class 1 as INJECTION. Its ONNX example returns logits. The adapter must
> validate the selected artifact's output contract and map the positive class
> correctly.
> * Convert two-class logits into probabilities using numerically stable
> softmax, where that is the selected export's contract. Do not apply softmax a
> second time to an already normalized probability output.
> * Return {{P(INJECTION)}} through the existing BOOLEAN probability
> representation. For example, a top-label response equivalent to BENIGN with
> probability 0.97 must yield an injection probability of 0.03, not 0.97.
> * Validate output dimensions, label mapping, finite numeric values, and
> probability bounds. Incompatible artifacts and malformed outputs are errors.
> * Preserve probability information until the common semantic threshold and
> uncertainty policy are applied. Do not collapse it to the model's default
> label first or introduce a conflicting hidden decision threshold.
> * Do not advertise this probability as an application-defined SCORE severity
> scale, or invent a calibrated confidence guarantee.
> BENIGN means that this evaluation did not detect injection. It does not
> establish general safety, authorization, or the absence of other threats.
> Actions such as block, quarantine, review, or continue belong to Camel
> routing policy; they are not additional model labels.
> h2. Tokenization and document handling
> Own tokenization, input tensors, special tokens, attention masks, padding,
> and model invocation inside the expert. Use tokenizer assets compatible with
> the pinned model and validate against reference inference.
> The Small v2 model card describes a 2,048-token window. Its published
> document evaluation uses overlapping windows with 64-token overlap and
> normalized Smooth-Max aggregation. A single truncated window does not
> reproduce that protocol.
> Define and document bounded handling of longer input. Either implement a
> documented document-scoring policy or reject inputs beyond the supported
> limit explicitly; never silently discard an unchecked suffix. If adopting the
> published aggregation, verify the exact formula and parameters against an
> authoritative implementation rather than guessing from its name. Clearly
> identify any alternative policy and its implications for threshold selection.
> For supported multi-window evaluation, test injection-bearing content beyond
> the first window and at window boundaries. Limits should bound both
> tokenization work and the number of inference windows. Empty, missing, and
> oversized input must have an explicit, tested outcome.
> h2. Expert configuration and lifecycle
> Keep model-specific configuration on the configured Wolf expert instance:
> model/tokenizer locations and revision, selected export, runtime options, and
> input/document limits. Named evaluations retain common semantic options such
> as state, threshold, and uncertainty policy.
> Support reproducible use of explicitly provisioned local artifacts without
> requiring downloads during normal route execution. Model weights should not
> be bundled into the Camel source repository or ordinary module artifact.
> Document acquisition, artifact identity, and compatible
> tokenizer/configuration files.
> Load and reuse model resources through Camel-managed lifecycle. Separate
> declaration/capability inspection from model loading and inference. Close
> sessions, tensors, and tokenizer/native resources on shutdown and startup
> failure. Support concurrent evaluations without cross-exchange state leakage,
> and bound concurrency and resource consumption using the selected runtime's
> facilities.
> Respect the semantic SPI's timeout, interruption, and shutdown contract.
> Verify the actual native-runtime cancellation behaviour; a timeout around a
> Java task must not be described as stopping native inference if it only stops
> waiting for it.
> Model-loading failures, invalid inputs, inference failures, malformed output,
> and cancellation must remain errors. They must not become a benign result or
> an injection probability of zero. Diagnostics should include enough
> provider/artifact identity to reproduce a problem without logging submitted
> text or credentials.
> h2. Integration with experts and Camel Catalog
> Register the expert through the discovery/lifecycle mechanism agreed in
> CAMEL-25310 and allow explicitly configured registry instances. Reuse its
> explicit/default/sole-expert resolution rules and its error on ambiguity. Two
> configured Wolf instances may use different artifacts or limits and must
> retain their separate identities.
> Publish generated static capability metadata for {{wolf-defender}}, linked to
> {{camel-wolf-defender}}, through the same authoritative definition used by
> runtime capabilities. Catalog inspection must work without downloading
> weights or initializing a native runtime. Static metadata describes the
> provider's contract; it must not claim that an arbitrary local artifact has
> been validated.
> Support the existing single-evaluation and batch SPI contracts. Framework
> grouping across experts remains the responsibility of CAMEL-25310. The
> initial implementation may use the existing sequential batch behaviour;
> optimized tensor batching is not required. Preserve named result keys and
> all-or-error publication.
> h2. Illustrative application declaration
> The following uses the syntax proposed in CAMEL-25310 and assumes a
> configured Wolf expert instance named {{security}}. It is illustrative,
> pending the final framework API and binding conventions.
> {noformat}
> - semantic:
> question:
> injection:
> expert: security
> type: boolean
> state: "${body}"
> threshold: "{{security.injection.threshold}}"
> uncertainty: "{{security.injection.uncertainty}}"
> uncertaintyPolicy: fail
> {noformat}
> Routes use {{ref:injection}} through the semantic language. Provide a
> runnable example showing expert construction/configuration, artifact
> provisioning, and routing on the resulting decision. Demonstrate that
> model-specific options remain outside the named evaluation. Explain how
> uncertainty and operational errors are handled by the route.
> h2. License and documentation
> The Small model card declares Apache License 2.0 and retains MIT terms for
> the upstream mmBERT-small portions. Preserve applicable notices for any
> redistributed material and document model provenance separately from
> runtime/tokenizer dependency licenses. Confirm the license of the exact
> selected artifacts as part of implementation.
> Document installation, supported runtime/platform combinations,
> configuration, fixed BOOLEAN semantics, probability mapping, limits, and
> error behaviour. Describe false-positive/false-negative limitations without
> making benchmark or latency claims for an untested Java implementation.
> h2. Validation and acceptance criteria
> # The new {{components/camel-ai/camel-wolf-defender}} module builds and
> integrates with Camel's normal generated metadata and documentation processes.
> # A named BOOLEAN evaluation without instructions works through
> {{camel-semantic}}, with both explicit expert selection and the framework's
> sole-expert discovery behaviour.
> # Unsupported result types, instructions, and criteria fail during
> declaration validation without running inference. Unknown/ambiguous expert
> selection follows CAMEL-25310.
> # Tests verify positive-class mapping, stable probability conversion,
> malformed/non-finite output rejection, and common threshold/uncertainty
> boundaries, including a high-confidence BENIGN result.
> # Input validation and document limits are explicit and tested. Oversized
> input is never silently partially screened; any supported windowed policy has
> boundary and late-window coverage.
> # Lifecycle and failure tests cover resource cleanup, concurrent use,
> cancellation/timeout behaviour, and errors remaining distinct from negative
> classifications.
> # Batch evaluation preserves keys, expert-instance identity, and complete
> result validation without exposing partial success.
> # Static capabilities are discoverable through Camel Catalog, agree with
> runtime reporting, and can be read without model/native-runtime
> initialization.
> # Normal unit tests use deterministic fixtures and require neither network
> access nor a large model download. Add an opt-in integration test using a
> pinned real artifact to verify tokenizer/inference parity against reference
> outputs with a documented numeric tolerance.
> # Provide a runnable local example and setup instructions, including the
> tested artifact/runtime, license/provenance, and representative benign,
> injection, and difficult-benign inputs. Keep adapter correctness evidence
> separate from claims about model detection quality.
> h2. Related issue and references
> * [CAMEL-25310: semantic experts, capabilities and catalog
> metadata|https://issues.apache.org/jira/browse/CAMEL-25310]
> * [Wolf-Defender Small model card and inference
> examples|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small]
> * [Wolf-Defender full
> model|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection]
> * [Wolf-Defender Small license
> file|https://huggingface.co/patronus-studio/wolf-defender-prompt-injection-small/blob/main/LICENSE]
> * [Wolf-Defender v2
> article|https://patronus.studio/en/posts/wolf-defender-v2-prompt-injection-detection-on-device]
> * [ONNX Runtime Java
> API|https://onnxruntime.ai/docs/get-started/with-java.html]
--
This message was sent by Atlassian Jira
(v8.20.10#820010)