[ 
https://issues.apache.org/jira/browse/IGNITE-28867?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18101922#comment-18101922
 ] 

Anton Vinogradov commented on IGNITE-28867:
-------------------------------------------

[Ignite PR Checker|https://ignite-pr-checker.is-a.dev/?pr=13335] verdict for PR 
13335 · RunAll build [9254327|https://ci2.ignite.apache.org/build/9254327] · 
147 suites ran, 0 reused

(!) *This run doesn't cover the PR fully:*
- 1 suite(s) have no reliable result (compilation error, timeout, crash)

(!) *1 broken suite(s)* (failed without a reliable run):
- Thin Client: Java: non-zero exit code · Number of tests 58 is 87% less than 
439 in build #17747 — ran 58 of master's 440 tests

(?) *No blockers found — but the run above can't prove the PR is clean.* 56 
pre-existing/flaky tests filtered out. Re-run once the above is sorted out.

⏳ _Auto re-run *#1* in progress — 1 broken suite(s) re-queued (attempt 1/2), *≈ 
settled by 05:17 MSK*. This comment updates when they settle._

> Add control.sh --ssl commands for TLS certificate hot redeployment
> ------------------------------------------------------------------
>
>                 Key: IGNITE-28867
>                 URL: https://issues.apache.org/jira/browse/IGNITE-28867
>             Project: Ignite
>          Issue Type: Task
>            Reporter: Mikhail Petrov
>            Assignee: Anton Vinogradov
>            Priority: Major
>              Labels: ise
>             Fix For: 2.19
>
>          Time Spent: 2.5h
>  Remaining Estimate: 0h
>
> Currently, Ignite node TLS certificates are loaded only once during node 
> startup and then used to establish secure connections between `node to node`, 
> `node to thin client`. As a result, updating a certificate requires 
> restarting the  Ignite node, causing unnecessary downtime and requiring 
> additional administrative effort.
> We need to implement hot redeployment of TLS certificates, allowing the 
> Ignite node to reload the certificates at runtime without requiring a node 
> restart. Existing TLS sessions should remain unaffected, while all new TLS 
> handshakes should use the updated credentials.
> Basic Steps
> 1. Introduce a mechanism to trigger certificate reloading (for example, 
> through `control.sh`).
> 2. Create a new `SslContext` instance by invoking `SslContextFactory`.
> 3. Replace the active `SslContext` used for TLS connection establishment. All 
> newly established connections must use the updated certificates, while 
> existing connections will continue to operate without interruption.
> h2. Design
> *Command.* {{control.sh --ssl reload}} is run against any node and reaches 
> every node of the cluster, servers and clients. {{--dry-run}} does the same 
> work but puts nothing in use, so the certificates on disk can be verified 
> before the cluster is touched.
> The command carries no certificate and no path: each node re-opens the key 
> and trust store files at the paths its own {{SslContextFactory}} is 
> configured with, and takes whatever is in them at that moment.
> *Registry instead of point-to-point calls.* A component registers itself in 
> {{GridInternalSubscriptionProcessor}} once it has set SSL up, so an empty 
> registry means the node does not use SSL. The command talks only to the 
> registry; a future SSL transport becomes reloadable by registering.
> {code:java}
> public interface SslContextReloadable {
>     boolean reloadSslContext() throws IgniteCheckedException;  // rebuild and 
> apply
>     boolean checkSslContext() throws IgniteCheckedException;   // rebuild only
>     default X509Certificate servedCertificate() { return null; }
> }
> {code}
> Registered today: communication, discovery, client connector, binary (TCP) 
> REST, HTTP REST.
> *Applying a new context.* {{AbstractSslContextFactory}} gains {{reload()}} 
> (rebuild and replace what {{create()}} hands out) and {{build()}} (rebuild 
> only, used by the dry run). Publishing into the factory is load-bearing: 
> outbound communication asks the factory per connection, and that is how it 
> picks the new certificates up. Inbound NIO transports get the context through 
> {{GridNioSslFilter}}, which creates an {{SSLEngine}} per session, so 
> established sessions are untouched.
> A factory that hands back the context already in use cannot be reloaded. This 
> is detected by identity rather than by type, and reported as {{NOT reloaded}} 
> instead of a false success.
> *Discovery.* The listening socket is now a plain {{ServerSocket}}, and every 
> accepted connection is wrapped with the context current at that moment. An 
> {{SSLServerSocket}} captures the context at bind time, which would have left 
> discovery reloadable for outbound connections only.
> *Check before applying.* For communication and discovery, a rebuilt context 
> must complete a mutually authenticated TLS handshake in memory before it 
> replaces the one in use; otherwise the node keeps its current certificates 
> and the command fails. Client-facing transports are not checked, because a 
> client legitimately holds a different key and trust store. Only a refused 
> handshake counts as a failure, so the check cannot become a new way of 
> blocking an emergency rotation. That handshake also yields the certificate 
> the node now serves, which the report prints.
> *HTTP REST (Jetty).* Its TLS comes from the user's Jetty XML rather than from 
> Ignite, so it goes through Jetty's own {{SslContextFactory}} instead of 
> {{Factory<SSLContext>}}. Jetty rebuilds in place and has no rollback: once 
> its {{reload()}} has failed, the connector serves no TLS at all until the 
> next successful one. The stores are therefore loaded on a throw-away factory 
> first, the context in use is pinned back if the rebuild fails regardless, and 
> the next attempt unpins it so that Jetty reads the files again. A 
> configuration that hands Jetty a ready-made context has nothing to re-read 
> and is reported as not reloaded.
> *Clients.* The command reaches nodes, not the applications connecting to 
> them, and those need no restart either. The JDBC and ODBC drivers build their 
> SSL context per physical connection, so the next reconnect uses whatever is 
> on disk. A thin client rebuilds its context from disk once a TLS handshake 
> has been refused: the trigger is the refusal itself rather than a change on 
> disk, so switching certificates can only happen where the client already 
> cannot connect, and never moves it ahead of the cluster.
> h2. Why the context is replaced rather than the key managers
> The alternative is to keep one {{SSLContext}} and swap the delegates inside 
> {{X509ExtendedKeyManager}}/{{X509ExtendedTrustManager}}. Elasticsearch had 
> exactly that and removed it in elastic/elasticsearch#30509 for FIPS 140-2, 
> where only SunJSSE managers may be used; Jetty and Kafka also rebuild the 
> context. Swapping managers additionally keeps the session cache and ticket 
> keys of the old context, so a resumed session is still served the old 
> certificate and trust cannot be revoked by a reload.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to