luigidemasi commented on code in PR #26813:
URL: https://github.com/apache/camel/pull/26813#discussion_r4114976325
##########
components/camel-ai/camel-semantic/src/main/docs/semantic-language.adoc:
##########
@@ -0,0 +1,358 @@
+= Semantic Evaluation Language
+:doctitle: Semantic Evaluation
+:shortname: semantic
+:artifactid: camel-semantic
+:description: Evaluate named questions about message content to produce
boolean decisions, categories and scores through provider adapters
+:since: 4.23
+:supportlevel: Preview
+:tabs-sync-option:
+
+*Since Camel {since}*
+
+The Semantic language evaluates named questions through a provider-independent
adapter.
+Boolean questions produce decisions, choice questions produce category
strings, and score
+questions produce numbers on an ordered rubric. Calls are synchronous and may
block while
+inference runs. Provider errors propagate through normal Camel error handling.
+
+== Dependencies and providers
+
+Add `org.apache.camel:camel-semantic` and a provider, such as
+xref:ROOT:typesafe-ai-component.adoc[TypeSafe AI] (`camel-typesafe-ai`), using
the same Camel version.
+The provider owns credentials, model selection, request timeout,
+concurrency limits and transport resources. TypeSafe AI uses
+`camel.component.typesafe-ai.*` settings even when a route contains no
TypeSafe AI endpoint.
+
+One advertised adapter is selected automatically. No provider, or multiple
distinct
+providers, is an error. Select an existing bean with
+`camel.language.semantic.adapter=myAdapter`, or a class using its fully
qualified name.
+Use plain names without `#bean:` or `#class:` prefixes. Registry lookup takes
precedence over
+class resolution. Class selection uses Camel's class resolver and injector.
+Created adapters are registered as `camelSemanticAdapter` and managed by the
context.
+A collision at that name is an error. Referenced beans retain their existing
lifecycle owner.
+
+Camel Main and Camel JBang bind `camel.language.semantic.adapter` and
+`camel.language.semantic.default-state` to the language. Embedded applications
can resolve
+`SemanticLanguage` from the context and call `setAdapter` and
`setDefaultState` before route setup.
+The catalog describes the generic `language` expression model; it does not
provide a dedicated
+semantic Spring Boot configuration class. Configure the language bean
explicitly in Spring Boot
+until generic-language starter configuration is available.
+
+== Named questions
+
+The xref:others:yaml-dsl.adoc[YAML DSL] supports declarations alongside
routes, including declarations after their use:
+
+[source,yaml]
+----
+- semantic:
+ question:
+ department:
+ type: choice
+ instructions: Which department should handle this message?
+ criteria:
+ billing: Invoices, payments and refunds
+ technical: Bugs, outages and technical problems
+ actionable:
+ type: boolean
+ instructions: Does this message contain an actionable request?
+ threshold: 0.8
+ uncertainty: 0.1
+ uncertaintyPolicy: fail
+ urgency:
+ type: score
+ instructions: Assess urgency
+ criteria:
+ - Routine request
+ - Time-sensitive request
+ - Immediate attention needed
+- route:
+ from:
+ uri: direct:tickets
+ steps:
+ - setProperty:
+ name: department
+ expression:
+ language:
+ language: semantic
+ expression: ref:department
+ - choice:
+ when:
+ - expression:
+ simple:
+ expression: "${exchangeProperty.department} == 'billing'"
+ steps:
+ - to: direct:billing
+ - expression:
+ simple:
+ expression: "${exchangeProperty.department} == 'technical'"
+ steps:
+ - to: direct:technical
+ otherwise:
+ steps:
+ - to: direct:review
+----
+
+Names are context-wide. Duplicate declarations across resources and unknown
references fail.
+Reloading a resource replaces its complete set of questions, including
removing declarations
+no longer present. Development-mode route reload also removes definitions from
deleted or renamed files
+before parsing replacements. This replacement does not make the surrounding
route reload transactional.
+Existing expressions resolve the current definition on their next evaluation.
+Loading declarations does not perform inference. Java applications can
register immutable
+`SemanticQuestion` definitions using
`SemanticQuestions.get(context).replace(source, questions)`.
+
+A question's optional `state` Simple expression overrides
+`camel.language.semantic.default-state`, whose default is `$\{body}`.
Selectors are compiled
+before evaluation; selected strings, maps and lists are passed as data and are
never evaluated recursively.
+A missing selected header fails instead of falling back to the body. Blank or
invalid selectors
+fail. The original message is preserved. `CamelSemanticResult` contains the
latest successful
+normalized result and is cleared before each evaluation, including one that
fails.
+
+State must be a string, map or list. For byte arrays or stream bodies,
explicitly select
+`$\{bodyAs(String)}`. Enable stream caching before evaluating a stream when
later processors
+also need to read it. Unsupported state types fail instead of being implicitly
converted.
+
+YAML declarations are provided by `camel-semantic` through the YAML
deserializer resolver SPI.
+Include both `camel-semantic` and `camel-yaml-dsl` when using them. The YAML
DSL does not pull
+in semantic evaluation, and Java applications using `camel-semantic` do not
pull in the YAML DSL.
+
+== Results and policy
+
+Only boolean questions can be predicates. A category string is never
implicitly a boolean.
+A probability-based boolean uses an inclusive threshold (default `0.5`). A
nonzero
+`uncertainty` defines an inclusive band around that threshold. `fail` (the
default) raises
+an error within the band; `non-match` returns false. An already-boolean
provider supports the
+default policy without inventing a probability. Additional policy requirements
must be
+supported by the provider. Choice results must name a declared criterion.
Scores range from
+zero to the last rubric index and may be fractional.
+
+Probabilities, provider confidence and selected values are separate. Missing
optional fields
+remain absent. Results can retain provider/model identity and usage metadata.
These fields
+do not imply equivalent quality or calibration when switching providers.
Timeouts, malformed
+answers and unsupported capabilities are errors, distinct from valid negative
decisions or
+unmatched Choice results.
+
+== EIP integration
+
+Use `language("semantic", "ref:name")` wherever an expression or boolean
predicate is accepted.
+
+[cols="1,3"]
+|===
+|EIP / integration point |Usage
+|Choice |Store a category with Set Property, then compare that property in
ordinary when predicates.
+|Filter |Use a boolean question as the filter predicate.
+|Validate |Use a boolean question; false follows normal validation failure
handling.
+|Set Header / Set Property |Store a category or score for explicit reuse in
later steps.
+|Aggregate correlation |Use a category as a correlation expression, retaining
tenant or case identifiers where needed.
+|Aggregate completion |Evaluate a boolean question against the accumulated
body, with a size or timeout limit. `eagerCheckCompletion` instead sees the
incoming exchange.
+|Recipient List |Map a category to a configured list of recipient URIs.
+|Routing Slip |Map a category to a predefined processing sequence.
+|Enrich |Map a category to a configured resource URI and use an ordinary
aggregation strategy.
+|Loop |Reevaluate a boolean question on updated state, with an explicit
iteration or time budget.
+|Sort |Score each item once, store the scores, then sort using a deterministic
comparator.
+|On Exception / retryWhile |Evaluate whether another attempt is worthwhile,
with an explicit retry budget checked before inference.
+|Contextual action validation |Use Validate after ordinary permission checks
and before executing the action.
+|===
+
+Destination mappings belong to trusted route configuration; provider output
should select
+known labels rather than supply unrestricted endpoint URIs. Expressions
reevaluate on each
+invocation. There is no implicit exchange-wide inference cache. Store results
explicitly
+when reuse is intended, and reevaluate after relevant input changes.
+
+=== Predicates and explicit reuse
+
+[source,java]
+----
+from("direct:actionable")
+ .filter().language("semantic", "ref:actionable")
+ .to("direct:accepted");
+
+from("direct:validate")
+ .validate().language("semantic", "ref:actionable")
+ .to("direct:valid");
+
+from("direct:tag")
+ .setHeader("department").language("semantic", "ref:department")
+ .setProperty("urgency").language("semantic", "ref:urgency")
+ .choice()
+ .when(header("department").isEqualTo("billing")).to("direct:billing")
+ .otherwise().to("direct:technical");
+----
+
+Storing the category before xref:eips:choice-eip.adoc[Choice] performs one
semantic evaluation each time execution
+reaches that Set Header or Set Property step. The branches compare the stored
result without
+calling the provider again. Place that step inside a loop when the decision
must be refreshed
+on each iteration. Nested decisions can use separate properties to retain
their own results.
+Boolean questions can also be used directly as ordinary when predicates. These
patterns use
+the existing EIP model and work with the Java, XML and YAML DSLs.
+
+=== Aggregation and destinations
+
+For xref:eips:aggregate-eip.adoc[Aggregate] and other Java APIs accepting an
`Expression`, use
+`new LanguageExpression("semantic", "ref:department")` from
`org.apache.camel.model.language`.
+For example, group messages with that expression and a
`GroupedBodyAggregationStrategy`,
+using `completionSize(10)` and `completionTimeout(5000)` to bound the group.
Include a trusted
+tenant or case identifier in the correlation key when messages must remain
isolated.
+
+A completion predicate can be obtained with
+`context.resolveLanguage("semantic").createPredicate("ref:actionable")`.
+The aggregation strategy must first put the accumulated conversation in the
selected state.
+Use `completionPredicate(predicate).completionSize(10)` to retain a
deterministic size limit.
+
+Classify once, then map the stored label to destinations supplied by the route
author:
+
+[source,java]
+----
+from("direct:dispatch")
+ .setProperty("department").language("semantic", "ref:department")
+ .process(exchange -> {
+ String department = exchange.getProperty("department", String.class);
+ exchange.getMessage().setHeader("recipients",
+ Map.of("billing", "direct:billing,direct:audit",
+ "technical", "direct:technical").get(department));
+ exchange.getMessage().setHeader("slip",
+ Map.of("billing", "direct:invoice,direct:archive",
+ "technical",
"direct:diagnose,direct:archive").get(department));
+ exchange.getMessage().setHeader("resource",
+ Map.of("billing", "direct:billingKnowledge",
+ "technical", "direct:technicalKnowledge").get(department));
+ })
+ .recipientList(header("recipients")).end()
+ .routingSlip(header("slip"))
+ .enrich().header("resource").aggregationStrategy((original, resource) -> {
+ original.getMessage().setHeader("knowledge",
resource.getMessage().getBody());
+ return original;
+ });
+----
+
+=== Changed state and sorting
+
+A loop must have a finite budget in addition to its semantic predicate.
Combine the predicate
+with a check of `CamelLoopIndex`, and let the loop body update the selected
state. Each new
+iteration evaluates the current content. For example, a predicate can first
check that
+`exchange.getProperty(Exchange.LOOP_INDEX, 0, Integer.class) < 5` and then
call the boolean
+semantic predicate, before `loopDoWhile` invokes a refinement processor.
+
+For sorting, evaluate `ref:urgency` once for each item and build a list
containing each item and
+its score. Use `.sort(body(), Comparator.comparingDouble(ScoredItem::score))`
with that stored
+score. The comparator must not call the provider: sorting can compare an item
multiple times
+and in an implementation-dependent order.
+
+=== Bounded semantic retry
+
+A boolean question can supply the
xref:manual::exception-advanced.adoc[retryWhile predicate]
+for an exception clause. Ask whether another attempt is worthwhile using the
current failure
+and selected request context. Restrict this to operations that are safe to
retry.
+
+IMPORTANT: `retryWhile` replaces the `maximumRedeliveries` decision. Check the
retry budget
+inside the predicate, before calling the provider. A provider that always
returns true must
+not cause unlimited retries or evaluations.
+
+Declare a boolean question named `retryable` with `state:
$\{exchangeProperty.retryState}`.
+In this example, the request body is a string. `onExceptionOccurred` prepares
the state before
+the retry predicate runs; `onRedelivery` runs later and is too late for this
purpose.
+Include only the failure details needed by the question, excluding credentials
and sensitive data.
+
+[source,java]
+----
+Predicate retryable =
context.resolveLanguage("semantic").createPredicate("ref:retryable");
+
+onException(IOException.class)
+ .onExceptionOccurred(exchange -> {
+ Exception failure = exchange.getProperty(Exchange.EXCEPTION_CAUGHT,
Exception.class);
+ exchange.setProperty("retryState", Map.of(
+ "request", exchange.getMessage().getBody(String.class),
+ "failureType", failure.getClass().getSimpleName(),
+ "attempt",
exchange.getMessage().getHeader(Exchange.REDELIVERY_COUNTER, Integer.class)));
+ })
+ .retryWhile(exchange ->
+ exchange.getMessage().getHeader(Exchange.REDELIVERY_COUNTER, 0,
Integer.class) <= 3
+ && retryable.matches(exchange))
+ .redeliveryDelay(1000)
+ .handled(true)
+ .to("direct:escalate");
+----
+
+The counter starts at one when deciding the first redelivery. This permits at
most three
+redeliveries after the initial attempt. Each failure refreshes the selected
state; redelivery
+restarts at the failed processor, not at the beginning of the route. A
negative decision or
+an exhausted budget sends the message to `direct:escalate` through normal
exception handling.
+
+Evaluation errors are distinct from a negative decision. If the retry
predicate throws,
+Camel reports the evaluation failure on the exchange; it does not
automatically execute
+the exception clause's escalation route. Arrange caller or supervising-route
handling for
+that failure. Bound provider timeouts as well as the retry count.
+
+=== Contextual action validation
+
+Use a semantic question for an additional check such as "Does this proposed
action serve
+the approved task?" after normal identity, permission and tenant checks have
succeeded.
+The semantic decision must not grant permissions that those checks denied.
+
+For example, declare this question alongside routes:
+
+[source,yaml]
+----
+- semantic:
+ question:
+ withinScope:
+ type: boolean
+ instructions: Does the proposed action serve the approved task?
+ state: ${body}
+ threshold: 0.8
+ uncertainty: 0.05
+ uncertaintyPolicy: fail
+----
+
+The selected state should contain the approved task and the proposed action.
Obtain the
+approved task from trusted application state; do not let the proposed action
redefine it.
Review Comment:
Addressed in `39261e079cb4`. The contextual-validation section now
explicitly identifies proposed-action text as untrusted input that may attempt
prompt injection, and states that a positive semantic decision must never
override application authorization checks. The catalog documentation mirror was
regenerated.
_AI-generated by Codex via /oss-address-review on behalf of
[luigidemasi](https://github.com/luigidemasi)._
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]