[
https://issues.apache.org/jira/browse/CAMEL-25499?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Luigi De Masi reassigned CAMEL-25499:
-------------------------------------
Assignee: Luigi De Masi
> camel-semantic: Design audit SPI, TUI history and observability hooks
> ---------------------------------------------------------------------
>
> Key: CAMEL-25499
> URL: https://issues.apache.org/jira/browse/CAMEL-25499
> Project: Camel
> Issue Type: Improvement
> Components: camel-ai, camel-jbang, camel-telemetry
> Reporter: Luigi De Masi
> Assignee: Luigi De Masi
> Priority: Major
> Attachments: camel-semantic-audit-tui.png,
> camel-semantic-audit-tui.svg
>
>
> h2. Purpose and status
> Design discussion for pluggable semantic auditing and an Audit view in Camel
> TUI, building on CAMEL-25491 / [PR
> #27602|https://github.com/apache/camel/pull/27602]. This issue captures a
> brainstorm and its open questions; it does not claim an implemented audit
> API, accepted upstream design or scheduled release.
> The implementation baseline inspected on 9 October 2026 is PR #27602 at
> {{0dcb5c2f15078331c0df2aad860d71c32e276c75}} (open at the time of
> inspection). It provides contract-driven Definitions, Experts and
> Relationships views, runtime metadata, explicit named-definition sample
> evaluation and direct expert calls, and bounded asynchronous
> connector/console evaluation.
> The starting observation is [Claus's CAMEL-25311
> comment|https://issues.apache.org/jira/browse/CAMEL-25311?focusedCommentId=18123470&page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#comment-18123470]:
> the expert's result, the policy applied and the action enforced by the
> caller must remain distinct. The same runtime evidence could support
> monitoring and an audit trail, with different consumers and delivery
> requirements.
> h2. Relationship to existing work
> * CAMEL-25491 / PR #27602 supplies the concrete runtime, connector and TUI
> baseline. Extend its Semantic tooling with a fourth Audit view.
> * CAMEL-25357 already proposes a recent-evaluation recorder, bounded buffer
> and console-backed view. Coordinate with and reuse/refine that recorder; do
> not introduce a competing recent-evaluation mechanism. This issue develops
> the pluggable backend contract, linked route-decision records, enablement
> semantics and shared observation lifecycle beyond that proposal.
> * CAMEL-23369 tracks broader GenAI observability. The common semantic hooks
> proposed here should feed that effort through existing Camel telemetry
> infrastructure, rather than create an independent OpenTelemetry SDK/exporter
> configuration.
> * CAMEL-25382 supplies expert-owned contracts, shared validation and typed
> result semantics.
> * CAMEL-25311 supplies the Wolf Defender example and the original
> observability discussion. The proposed audit API must remain
> provider-independent and also support ordinary business
> classification/scoring.
> h2. Working requirements from the brainstorm
> * Audit capture runs in the integration and continues without a connected TUI.
> * Audit backends are pluggable. Writing and querying are separate
> capabilities so a write-only destination does not have to implement browsing.
> * Auditing and OpenTelemetry can be enabled independently. Audit enabled with
> OpenTelemetry disabled is a supported combination.
> * A master audit setting supplies the default. An explicit per-expert value
> overrides it in either direction. Experts without an override inherit the
> master value.
> * Overrides identify configured expert instances by their registry bean name,
> not by provider class or static annotation identity.
> * Disabling audit capture does not disable inference, erase retained records
> or prevent reading existing history.
> * The UI is a screen inside the existing Camel TUI. No webpage is proposed.
> * Configuration placement remains open: application properties, semantic DSL,
> or both. The examples below describe alternatives, not existing supported
> syntax.
> h2. Record model: evaluation evidence and route decisions
> Use immutable, append-only records with two principal categories:
> * *Evaluation*: which expert operation ran, its typed result and documented
> meaning, optional probability/confidence, duration and execution status.
> Camel captures this at the semantic runtime boundary.
> * *Decision*: the route/policy's explicit action (for example allow, block,
> review, quarantine or continue), policy identity and reason, referencing the
> evaluation record(s) that supported it.
> An expert returning {{true}} for an injection operation does not prove that
> the route blocked a request. A permitted tool call does not establish
> successful tool execution. A policy can use several evaluations or make a
> decision without one; keep those relationships explicit. A later decision
> references earlier evaluation records instead of mutating them.
> Provide a route-facing API usable through normal Camel Java and YAML/XML bean
> invocation to record application decisions. Do not infer business policy
> meaning from an arbitrary Choice/Switch branch. The API shape and how a route
> obtains the relevant invocation references remain design questions.
> h3. Table fields and their ownership
> ||Field||Meaning / source||
> |Timestamp|Event occurrence time, stored in UTC; display timezone is a
> presentation choice.|
> |Decision / action|Explicit route/policy action; absent for an evaluation
> alone.|
> |Category|Evaluation or decision. Lifecycle/request observations may require
> an additional discriminator.|
> |Operation|Application operation such as {{tools/call}}; keep the expert
> operation, such as {{injection}}, separately.|
> |Target|Logical tool, resource or business target supplied by the
> application. An evaluation row can display its definition name as its target.|
> |Namespace|Optional application/tenant namespace from trusted application
> context or configuration; do not infer it from arbitrary inbound headers.|
> |Reason code|Stable runtime or application-policy code; keep explanatory text
> separate. A code such as {{governance_no_match}} cannot determine allow/block
> by itself.|
> |Correlation ID|Application correlation identifier. Keep Camel exchange
> identifiers and tracing identifiers as distinct fields.|
> |Details|Structured evidence and linked records, shown in the TUI inspector.|
> Additional detail fields:
> * Schema version, unique event ID, invocation ID, optional batch/attempt
> identifiers, and links to supporting evaluation events.
> * Application/Camel context, route and processor IDs where available.
> * Named definition when present, expert bean reference and static
> provider/operation identity.
> * Origin: route execution, definition sample, or direct expert call. Do not
> claim a particular client identity unless it was propagated and is known.
> * Typed result and its meaning, optional scalar/label probabilities and
> confidence, with expert-defined semantics. Missing values stay absent.
> * Allowlisted provider/model/revision metadata and supplied policy
> ID/version/rule.
> * Duration, execution status and sanitized failure code. Failures, invalid
> results and uncertain outcomes retain their declared semantics; they do not
> become benign answers.
> Capture the permitted result semantics needed to interpret historical
> records. In particular, PR #27602's {{scoreLevelsParameter}} and submitted
> level descriptions must not be replaced by today's edited definition when
> displaying an older score. Sensitive descriptions can be omitted/redacted.
> h2. Proposed runtime and backend boundary
> Illustrative names and signatures, subject to discussion:
> {code:java}
> public interface SemanticAuditSink extends Service {
> void append(SemanticAuditRecord record) throws Exception;
> }
> public interface SemanticAuditReader {
> SemanticAuditPage query(SemanticAuditQuery query);
> Optional<SemanticAuditRecord> get(String eventId);
> }
> {code}
> Register configured backends as ordinary Camel beans. A context-managed audit
> service handles lifecycle and dispatch. Multiple sinks can receive the same
> record; configure the reader used by the TUI explicitly. The sink
> acknowledgment/durability contract, threading requirements and pagination
> contract need definition before implementation.
> Candidate initial backends:
> * A bounded in-memory recorder/reader for recent TUI history, coordinated
> with CAMEL-25357. Report capacity and evictions; this is not durable storage.
> * Structured JSON logging to integrate with existing collection
> infrastructure.
> * The SPI permits database, messaging and external audit-service integrations
> without provider-specific changes. Which persistent backend belongs in the
> first implementation remains open.
> h3. Capture points in PR #27602
> * {{SemanticLanguage}} has named-expression/batch execution and a separate
> {{evaluate(SemanticEvaluation, Object)}} direct path. Cover both, including
> common validation failures; instrumenting only {{SemanticEvaluateConsole}}
> would miss normal routes.
> * Preserve input validation before batch provider calls. Distinguish
> completed members, failed groups and members never executed if a later group
> fails. Do not split a provider's optimized batch into individual calls merely
> to audit it.
> * Give repeated expressions, nested evaluations and redelivery attempts
> distinct invocation identities. Export retries retain the same event ID.
> * {{SemanticEvaluateConsole}} supplies sample/direct origin and request
> identity. Completion must not be recorded twice merely because both the
> console and language observe it.
> * Startup validation and metadata discovery do not perform inference or
> produce fabricated evaluation records.
> * Console timeout/disconnect requests cancellation but may not stop
> native/remote inference. Record the request outcome separately from observed
> provider completion, linked by invocation ID. Define late-completion handling
> explicitly.
> h3. Data handling and delivery
> * Create safe immutable snapshots before asynchronous dispatch. Do not retain
> the live {{Exchange}}, mutable provider objects, message body, input text,
> headers, credentials or unfiltered exception messages.
> * Apply allowlisting/redaction and size bounds before records reach any sink,
> query endpoint or telemetry consumer. Provider metadata is not a blanket
> permission to persist arbitrary data.
> * Proposed initial delivery is bounded and asynchronous, with visible backend
> errors, dropped-record counters and lifecycle-aware shutdown. Observer
> failures must not rewrite expert results or route policy actions.
> * Document whether a successful append means queued, written or durably
> committed. A memory record does not prove another sink persisted it. Retries
> and deduplication must use stable event IDs.
> * Guaranteed recording before a protected side effect would require a
> separate explicit delivery mode. It is not promised by this brainstorm's
> asynchronous baseline. Crash behavior, partial delivery, backpressure and
> retention are open design details.
> h2. Audit enablement semantics
> The master setting is a default that explicit expert settings can override,
> not an unconditional disable-all control:
> ||Master||Expert override||Effective evaluation audit||
> |false|omitted|disabled|
> |false|true|enabled|
> |false|false|disabled|
> |true|omitted|enabled|
> |true|true|enabled|
> |true|false|disabled|
> For example, master enabled + {{security=true}} + {{decisions=false}} audits
> the security expert but not the decision expert. Setting the master to false
> still audits {{security}} because its explicit true wins. Removing that
> override restores inheritance.
> Use cases include retaining guardrail evidence without all routine
> classifications, reducing volume from frequently invoked experts, and
> enabling targeted investigation for one configured instance. Per-expert
> filtering concerns evaluation records. The treatment of explicitly recorded
> route decisions, particularly decisions based on multiple experts, must be
> specified independently; do not arbitrarily inherit one supporting expert's
> filter.
> h3. Open point: configuration placement
> Both properties and the semantic DSL are valid candidates. No choice was made
> in the brainstorm.
> * Properties support deployment-specific settings and centralized operator
> overrides, but separate audit choices from the route definitions they affect.
> * DSL settings make the relationship visible and versioned with declarations,
> but couple operational capture choices to route resources and their reload
> behavior.
> * Supporting both needs a documented precedence and reload model. Avoid
> multiple conflicting authorities.
> Properties illustration (*proposed audit keys; not available in PR #27602*):
> {code:properties}
> camel.semantic.audit.enabled=true
> camel.semantic.audit.experts.security.enabled=true
> camel.semantic.audit.experts.decisions.enabled=false
> # Independent existing Camel OpenTelemetry2 setting
> camel.opentelemetry2.enabled=false
> {code}
> Alternative DSL illustration (*proposed syntax; not available in PR #27602*):
> {code:yaml}
> - semantic:
> audit:
> enabled: true
> experts:
> security:
> enabled: true
> decisions:
> enabled: false
> evaluation:
> screenPrompt:
> expert: security
> operation: injection
> state: "${body}"
> {code}
> The audit settings belong to the semantic framework, not provider bean
> properties or the static {{@SemanticExpert}} contract. The default value when
> nothing is configured, whether changes are live, and validation of
> unknown/late-registered expert references remain open.
> h2. Shared foundation for observability
> Use provider-independent semantic lifecycle hooks and a shared safe metadata
> model. Audit recording and OpenTelemetry instrumentation consume those hooks
> independently:
> {noformat}
> Semantic evaluation start / completion + explicit route decision
> |-- audit consumer --> configured sinks --> reader --> TUI
> `-- telemetry consumer --> Camel telemetry / configured SDK exporters
> {noformat}
> Completed records can supply structured events and counters. Full tracing
> also needs start/completion hooks and parent context propagation, including
> across the bounded executor boundaries introduced by #27602. Hook names and
> interfaces remain illustrative; keep OpenTelemetry types/dependencies outside
> the generic semantic API.
> * Support audit on / telemetry off, audit off / telemetry on, and both on.
> * The agreed audit master and expert overrides govern audit capture.
> Telemetry uses its own configuration; trace sampling must not suppress audit
> records.
> * Reuse Camel telemetry / {{camel-opentelemetry2}} integration and coordinate
> with CAMEL-23369. Do not introduce another independently configured SDK or
> OTLP client.
> * Preserve expert results and caller actions separately in telemetry.
> Successful detection of injection is not an inference failure.
> * Keep high-cardinality correlation IDs out of metric labels. Preserve them
> in records/traces where appropriate.
> * Apply guardrail conventions only to relevant operations, not all business
> classifications/scores.
> [OpenTelemetry guardrail PR
> #427|https://github.com/open-telemetry/semantic-conventions-genai/pull/427]
> is still open at inspected head {{7b8e842f84db31d0c8784ff26390d0eba9414593}}.
> Its [current event
> definition|https://github.com/open-telemetry/semantic-conventions-genai/blob/7b8e842f84db31d0c8784ff26390d0eba9414593/model/gen-ai/events.yaml]
> uses {{gen_ai.guardrail.result}}; names have evolved since the original
> comment. Keep the mapping versioned and isolated from the audit schema
> instead of treating draft names as a stable Camel API.
> h2. TUI integration and mockup
> Add *Audit* alongside *Definitions*, *Experts* and *Relationships* in the
> Semantic screen from #27602. The attached mockup follows those screens'
> orange focus accents, blue selection, terminal frame and keyboard hints. It
> uses illustrative data and is not an implemented screen.
> !camel-semantic-audit-tui.png|width=1200!
> Attachments: [PNG preview|^camel-semantic-audit-tui.png] and [editable
> SVG|^camel-semantic-audit-tui.svg].
> * A chronological table shows the fields above, with a selected-record
> inspector separating route decision from linked evaluation evidence.
> * Filter by time, category, action, expert, route, namespace and correlation
> ID. Use bounded cursor-based queries and preserve selected event identity
> during refresh.
> * Show effective audit master/expert settings, telemetry state, selected
> backend, retention/evictions and delivery problems. Disabled capture and
> empty history must be distinguishable.
> * Use text labels as well as colors. Adapt columns to terminal width, moving
> secondary fields into details when needed.
> * Query through {{SemanticTab}} / a focused {{SemanticAuditView}},
> {{MonitorContext.executeIndependentAction}}, a new {{semantic-audit}}
> connector action and a read-only {{SemanticAuditConsole}} backed by the
> configured reader.
> * Keep audit query state separate from the definition metadata snapshot.
> Browsing history must not require that an old definition still exists or run
> an expert.
> * Extend {{LocalCliConnector}}'s request-specific asynchronous dispatch for
> potentially slow audit queries, with bounds independent of inference workers.
> Preserve its stop/disconnect behavior and avoid blocking normal
> polling/control actions.
> * Reuse {{SemanticResultView}} formatting helpers using captured semantics;
> remove playground-specific draft/retry instructions from historical records.
> * Extend the existing TUI MCP inspection surface with visible audit rows,
> filters and selected details. This is inspection, not inference replay.
> h2. Proposed validation and completion criteria
> Before implementation, settle API/data ownership, configuration placement,
> delivery/acknowledgment semantics and coordination with CAMEL-25357 /
> CAMEL-23369.
> An implementation should then demonstrate:
> # Named, direct, nested and batch evaluations produce correctly correlated
> records without changing provider invocation count, policy interpretation or
> normal Camel error handling.
> # Input/result failures, provider errors, redelivery, timeouts, cancellation
> requests and late completion are represented accurately without stale or
> duplicate outcomes.
> # All six master/override combinations above work by expert instance, and
> audit-only / telemetry-only operation works independently.
> # Explicit decisions can reference multiple evaluations; their recording
> policy is documented separately from expert filtering.
> # No submitted content, secrets, mutable exchange state or unfiltered
> provider metadata leaks through storage, querying, TUI or telemetry.
> # Bounded capacity, concurrency, sink failure, partial delivery, retry
> identities and shutdown/restart behavior match the documented contract. No
> durability claim exceeds backend guarantees.
> # The TUI remains responsive during slow inference/backend queries, handles
> paging/eviction/disconnection and shows historical result semantics correctly.
> # Relevant documentation, console descriptors/catalog mirrors and tests are
> updated without expanding the static expert contract into an audit-backend
> contract.
> _Generated by OpenAI Codex on behalf of @luigidemasi._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)