Hi Prithvi, Thanks for bringing this topic up - it is indeed quite important that we, as a community, work on making this feature ready for tricky and highly-regulated environments. Here are my thoughts on the questions you've posed in your email.
> 1. Confirm that external listeners are the supported production audit path, and that polaris_schema.events is not. I don't think I agree with this. As the code implies, the `PolarisPersistenceEventListener` is just one of a handful of implementations Polaris strives to support (this is what is supposed to be writing to `polaris_schema.events`). It is the sole choice of the operator/implementor to use the in-persistence database and/or an external service to hold the events. > 2. Whether authentication failures and authorization allow/deny should become first-class events. Success-only coverage is not enough for an audit trail. The community discussed this earlier during the feature's implementation. The community made a soft decision not to support "deny" events because there wasn't much of a use case at that time. I'm sure we can reconsider if you have a strong use case you'd like to present. The important part here is the "so what?" - i.e., why would anyone want to log every single unauthenticated/unauthorized/unsuccessful call made to the server? How does generating an event on every potential request work in case of a DDoS attack on the Polaris server? > 3. Whether Polaris should document a production profile that does not silently drop events, or fail closed when the sink is down. This was another topic of discussion during the feature's implementation. The question really comes down to whether we ok with causing a server degradation/outage if the audit backend is not functioning. The earlier decision was that events should not interfere with query processing - especially since waiting for another HTTP roundtrip to an external system before returning the request can cause heavy negative side effects on request latency and/or availability. The event system was deliberately rewritten using an async mechanism as a result. However, I can still see a use case here where some users in very secure environments would be willing to sacrifice latency and/or availability for guaranteed auditing - I would be glad to help review such a rewrite if the community is supportive of this. > 4. Whether phase 1 is docs only (recommended listener set, delivery semantics, how to investigate a 403 in the sink), with coverage and delivery changes as follow-ups. Not sure what you mean by this. If you are asking whether this feature set is incomplete, the answer is yes. Many contributors have helped with certain parts of the feature set - and we will still require many more contributors and contributions. We welcome all contributions on this feature set. I don't believe anyone has claimed they are currently working on all these tasks, so I invite you to contribute to the pieces you see as non-controversial while we gather community consensus on this thread. I also see that you started a GH issue for this. I would recommend redirecting people to this ML thread instead of interacting on the GH issue, as tracking two separate discussions on the same topic would be very tough. Best, Adnan Hemani On Mon, Sep 21, 2026 at 2:25 PM Yufei Gu <[email protected]> wrote: > Hi Prithvi, > Thanks for outlining the gaps. I'm not sure we need stronger consistency > guarantees between request processing and audit persistence, or fail-closed > behavior when the sink is unavailable, at this stage. It would help to > understand the concrete requirements and availability tradeoffs before > committing to those guarantees. Documenting the current delivery semantics > sounds like a useful first step. > > Expanding authN/authZ event coverage is something we can move forward with. > Authentication failures and authorization allow/deny decisions would be > useful first-class events, including identity and request context where > available. We can improve that coverage independently of the delivery > guarantees discussion. > > Yufei > > > On Sun, Sep 20, 2026 at 12:23 PM Prithvi S <[email protected]> > wrote: > > > Hi all, > > > > I opened https://github.com/apache/polaris/issues/5561 to discuss a > > supported audit trail for Polaris operators. > > > > The event pipeline is already the right foundation: listeners since 1.0, > > JDBC events + CloudWatch in 1.2.0 (preview), multi-listener in 1.5.0, the > > async executor in 1.6.0, and Kafka plus OpenTelemetry listeners in 1.7.0. > > Events can carry principal, event type, timestamp, request id, and > > catalog/namespace/table identifiers, and DefaultEventSanitizer keeps > > credentials out of the default payload. > > > > That is not yet an operator-facing audit story. HTTP access logs and the > > default server log format do not reliably include the principal. Listener > > delivery is best-effort (bounded backlog drops; the JDBC buffer is > > in-memory, async, and outside the request transaction). Authentication > > failures never reach the event delegator. Authorization denials are INFO > > log lines, not PolarisEventTypes. polaris_schema.events has no query API, > > secondary indexes, purge, or Polaris ACL, and NoSQL does not implement > > writeEvents. > > > > This is not a request for Polaris to become a SIEM, add a Management API, > > or build an in-process query engine. > > > > Questions for this thread: > > > > 1. Confirm that external listeners are the supported production audit > path, > > and that polaris_schema.events is not. > > 2. Whether authentication failures and authorization allow/deny should > > become first-class events. Success-only coverage is not enough for an > audit > > trail. > > 3. Whether Polaris should document a production profile that does not > > silently drop events, or fail closed when the sink is down. > > 4. Whether phase 1 is docs only (recommended listener set, delivery > > semantics, how to investigate a 403 in the sink), with coverage and > > delivery changes as follow-ups. > > > > Cheers, > > Prithvi S > > >
