Hi Zihan, Thank you — this is an extremely thorough review, and I've verified every claim against the tree. To answer your opening question first: the review commit faff2de is the current HEAD, so nothing here is stale. I reproduced the structure of all thirteen findings in source, and your mechanism descriptions hold line-for-line.
F1 is the root cause and it's real. connect_timeout was intended to bound only the TCP handshake — the doc comment on ConnectionOptions.connect_timeout says exactly that — and per-RPC deadlines were simply never wired up. That intent doesn't survive contact with reality: a wedged read is unbounded, auto-reconnect can't fire (it only reacts to Error::Thrift), and drop hangs forever. I consider this a defect, not a contract. Plan: add a socket_timeout to ConnectionOptions applied as set_read_timeout/set_write_timeout on the stream in Connection::open, set a deadline before the TLS handshake (which will make the currently-dead WouldBlock arm reachable — good catch on the consequences), and make sure it covers the whole lifetime including sweeps and the pool-closed teardown path. F2 — your read is right. The RPC-on-drop was deliberate, to match the C#/Node SDKs' best-effort close, but with F1 present it is non-recoverable rather than best-effort, and I hadn't fully appreciated that a wedged destructor also strands the other seven sessions in a pool close and hangs an unwind. Once socket_timeout exists, the drop path becomes bounded, which I think is the right fix rather than dropping the transport without a closeSession — but I'm open to the alternative if you have reasons. F4–F7 — all accepted, all real. authenticate running outside the endpoint loop is an oversight (reconnect does the right thing; open should too). F5's budget-less growth failure and F6's lost-Condvar-notification are both legitimate; I'll fix the accounting so live is only touched under the lock. On F7: the acquire-side eviction loop is defensive — with enable_auto_reconnect = false it's the only place a dead session gets caught on the hot path — but you're right that it's hard to reach in production today; I'd rather make is_open reflect socket state than lean on it. F8–F13 — accepted. F8's hold-the-slot-while-reconnecting is worth fixing; F9 contradicts its own doc comment; F10's frame cap matches Go/C# so it's not a regression, but I'll look at making it configurable and draining the body before rejecting; F11/F12 (host-string equality, queue-position-not-recency) I'll fix — preferring the newest hint is strictly better given RedirectCache already tracks seq. F13's Duration::MAX panic is a fair catch; I'll saturate instead of overflowing. The two smaller observations: the fixed TTL/max_entries should become SessionConfig knobs — agreed, and the reason they're fixed is simply that they predate the config plumbing. TableSessionPool missing acquire_for_device is an omission, not deliberate; I'll add it so table-mode hints have a consumer. The "checked, nothing found" section was as valuable as the findings — thank you for recording it. I'll open the F1–F7 items as GitHub issues in the next few days and start on the socket timeout. Best regards, Xuan Wang Zh D <[email protected]> 于2026年8月4日周二 23:14写道: > Hi Xuan, all, > > A review of the session pool, connection lifecycle and redirect > handling in apache/iotdb-client-rust, at faff2de. > > I previously reviewed the Node.js client along the same axes -- pool > acquire/release, connection lifecycle, redirect handling -- and > applied the same method to the Rust implementation. Everything below > was read in the source and, where noted, reproduced locally against > a throwaway fake Thrift server (handshake answered normally, then > silent) or a blackhole listener. Where I only traced statically, I > say so. > > > There is no client-side time bound after the TCP handshake > ---------------------------------------------------------- > > This is the one finding I would put first. F1, F2 and F3 below are > three fix points for it, kept separate because they are fixed in > different places, but they are one root cause. > > connect_stream (connection/mod.rs:246-261) applies connect_timeout > to TcpStream::connect_timeout only, which covers the TCP > handshake. After that, open (:180-181) sets only set_nodelay(true). > There is no set_read_timeout / set_write_timeout anywhere in the > crate and no socket2 dependency, so every Thrift read after the > handshake is an unbounded blocking read. ConnectionOptions > (:87-106) has no field for one. SessionConfig::query_timeout_ms > (session.rs:61-62) is a request-body field carried in > TSExecuteStatementReq (:840) and TSFetchResultsReq (:484) -- it is > enforced server-side, gives no client-side bound, and does not > cover openSession or closeSession. Rust's TcpStream does not enable > SO_KEEPALIVE by default either, so if the peer's kernel never sends > FIN/RST the block is unbounded in the literal sense. > > F1. Ordinary RPCs. Against a listener that completes the TCP > handshake and then never replies -- equivalent to a long JVM GC > pause, a stateful firewall silently dropping return traffic, or an > LB that accepts but does not forward: > > - Session::open() with connect_timeout = 200ms was still blocked > after 12s; the block is in the openSession read (session.rs:278), > past the connect timeout's scope. > - With a fake server that answers openSession + requestStatementId > normally and then goes silent, execute_non_query blocked for >8s > with enable_auto_reconnect = true and max_reconnect_attempts = 3. > The reconnect path never fired, because a wedged read produces no > Error::Thrift and with_retry (session.rs:349-357) only reacts to > errors. So auto-reconnect does not help against the one failure it > looks like it should cover. > - Control: with the port closed, the same call returned in 176us. > The hang is specific to "handshake succeeded, peer silent". > > Reachability is the default configuration through the public API: > Session::open, any RPC, TableSession::build, SessionPool::new (which > calls open_session() at pool.rs:150 inside the constructor when > min_size > 0), SessionPool::acquire. There is no escape hatch for a > user outside the library -- Connection's only public constructor is > open, and the TcpStream is never exposed. > > Question. Is the intended contract that connect_timeout bounds only > the TCP handshake and that per-RPC deadlines are out of scope for > now, or was a socket-level read timeout intended and simply not > wired up? Asking because several of the findings below are > amplifications of this one and would collapse if a socket_timeout > existed. > > F2. Drop, where the caller has no recourse. Three destructors send > RPCs that wait for a response: Drop for Session (session.rs:847) -> > close() (:822) -> close_session (:825); Drop for SessionPool > (pool.rs:392) -> close() (:307) -> serial entry.session.close() for > every idle session (:314-316); Drop for SessionDataSet > (dataset.rs:213) -> close() (:202) -> Session::close_query > (session.rs:500) -> close_operation. Drop for PooledSession > (pool.rs:427-433) is indirect: release() only issues RPCs on the > pool-closed branch (:370) or when the sweep expires entries > (:373-377); the normal return path (:372) just pushes to the idle > queue. > > A destructor cannot return an error, cannot be cancelled, and > cannot be given a timeout by the caller. The failures are also > swallowed (session.rs:826, pool.rs:315, :370, :376 are all let _ = > or log::debug!), so a wedged teardown produces neither an error nor > a warning. Reproduced against the same silent server: > > 1. Session::open() succeeds -- the server log shows ["openSession", > "requestStatementId", "closeSession"], i.e. closeSession was > written -- then drop(s) blocked >12s and will not return. > 2. SessionPool::new(min_size = 1) then drop(pool): blocked >12s at > pool.rs:314-316. With max_size = 8 idle sessions, the first one > blocks, the other seven are never attempted, and notify_all() at > :317 is never reached. > 3. pool.close() followed by a guard going out of scope: blocked > over 12s on the release() closed-branch at :370. > 4. If a panic unwind passes through a bare Session or a > SessionPool, the unwind never completes. > > A secondary effect on waiters, which is latency only: in close(), > closed = true is set inside the lock (:310) and the lock is > released at :312 before the blocking loop, so waiters are not stuck > on the mutex -- good. But notify_all() is at :317, after the > blocking closes. With two or more waiters parked in wait_timeout > (:254), the one not woken by the earlier notify_one() in release() > (:379) sleeps out its full acquire_timeout. Measured with > acquire_timeout = 4s: one waiter returned in 305ms, the other in > 4.004s. Both got the correct "session pool is closed" error. > > Question. Would you consider a teardown path that drops the > transport without waiting for closeSession, or is the best-effort > RPC-on-drop deliberate (matching the C#/Node SDKs)? It reads > deliberate -- I mostly want to flag that with F1 present it is > non-recoverable rather than best-effort. > > F3. TLS, where a one-line misconfiguration hangs instead of > erroring. Severity is lower because tls is a non-default feature. > At connection/mod.rs:185 the stream handed to tls_handshake is a > blocking socket with no read timeout, so connector.connect(domain, > stream) at :310 has no time bound. The comment at :312 ("Blocking > sockets never yield the mid-handshake variant") confirms the > blocking socket is intentional, which is precisely why a deadline > would have to be set before :310. The doc comment at :176-178 says > the TCP connect is bounded by connect_timeout and does not claim > the TLS step is, so this is an undocumented unbounded step rather > than a contradiction. > > Reproduced with use_ssl = true, connect_timeout = 200ms, pointed at > a non-TLS IoTDB RPC port. The client sends ClientHello and waits > for ServerHello; the server's framed transport reads the first four > bytes 16 03 03 00 as a frame length of 369,296,128 and waits for a > 369MB body. Neither side closes. Session::open() was still blocked > after 15s. Control: the same silent peer with tls: None returned in > 200.5us, so the non-TLS path is genuinely bounded and the TLS path > is not. The multi-endpoint failover loop in Session::open > (session.rs:238-247) never reaches the second endpoint, and > reconnect's max_reconnect_attempts / retry_interval are equally > never reached. > > The existing TLS tests do not cover this shape: > tls_dispatch_sends_client_hello (connection/mod.rs:621-643) and > tls_option_selects_stack_against_plain_endpoint (:646-676) both use > listeners that close the connection, so the client escapes on EOF > with Err(Error::Tls(_)). One consequence for whoever fixes this: > once a read timeout exists, the currently dead > HandshakeError::WouldBlock arm at :313-315 becomes reachable and > the comment at :312 stops being true. > > > Other findings, in brief > ------------------------ > > Each of these was reproduced or traced the same way; I have kept > them to a line each rather than expanding the mechanism. Happy to > open any of them as separate GitHub issues if that is easier to > track. > > F4 (medium) session.rs:238-252, contrast :313-337 -- Session::open > fails over only at the TCP layer: authenticate runs outside the > endpoint loop, so a handshake failure returns instead of trying the > remaining nodes. reconnect() does the opposite. With a 3-node > cluster where one node accepts TCP but fails openSession, roughly > one open() in three fails while 2 of 3 nodes are healthy. > > F5 (medium) pool.rs:235-242, :335-343 -- acquire() spends none of > its acquire_timeout budget when the growth branch fails (measured > 198us against a 5s budget; 4136 of 4800 acquires failed instantly > under stress while the pool had sessions circulating), and a failed > USE replay in hand_out discards the remaining idle candidates. > > F6 (medium, latency only) pool.rs:238-239, :339-340, :355-356 -- > live is decremented outside the state lock at five sites; at the > three above, the decremented value is exactly the predicate a parked > waiter re-evaluates, so a Condvar notification can be lost. The > comment at :121-122 no longer matches the code. Measured stalls > track acquire_timeout exactly (300ms -> 305ms, 1000ms -> 1.005s). > > F7 (medium, gated on enable_auto_reconnect = false) > session.rs:368-370, pool.rs:223-228 -- is_open() is > connection.is_some(), not socket state. With auto-reconnect off, a > dead session is reused indefinitely under load. Separately, I could > not construct a production path where the acquire-side eviction > loop evicts anything; is it defensive, or have I missed a writer? > > F8 (low-medium) session.rs:299-340, sleep at :315 -- reconnect() > sleeps and walks all endpoints while still holding its pool slot; > worst case with defaults is ~92s for three endpoints. Concurrent > callers get "pool exhausted", which points away from the cause. > > F9 (low) connection/mod.rs:246-256 -- connect_timeout is applied > per resolved address, not per endpoint, contradicting both doc > comments; a dual-address host multiplies every bound (measured > 2.0011s for two blackhole addresses at 1s each). > > F10 (low/informational) connection/mod.rs:231 -- the frame cap is > thrift-rs's default 16,384,000 bytes, which a conforming IoTDB > server can legitimately exceed. Go and C# use the same number, so > this is not a Rust-specific regression; the part worth attention is > that rejection happens before the body is drained, leaving the > connection desynchronized, and fetch_results does not go through > with_retry. > > F11 (low) pool.rs:284-286 -- redirect endpoint matching is exact > host-string equality with no normalization, so hostname-vs-IP > spelling silently disables the optimization on multi-DataNode > clusters. No log line, so it is invisible from outside. > > F12 (low) pool.rs:279-282 -- when idle sessions hold conflicting > hints for one device, the winner is queue position, not recency. > Within documented behaviour (:266-268 says "any idle session"), but > RedirectCache already tracks seq/inserted -- was preferring the > newest hint considered? > > F13 (low) pool.rs:207 -- acquire_timeout = Duration::MAX panics on > Instant + Duration overflow before the lock is taken. Narrow (100 > years and 1000 years both work), and it matters only because there > is no "never time out" option and Duration::MAX is the natural way > to ask for one. > > Two smaller observations: RedirectCache's TTL and max_entries are > fixed at construction with no SessionConfig knob, which becomes > measurable only on large batch inserts; and TableSessionPool > exposes acquire() but not acquire_for_device(), so table-mode users > accumulate hints no pool API can consume -- deliberate? > > > What this implementation gets right > ----------------------------------- > > These are the specific failure modes that bit the Node.js client, > absent here by construction rather than by discipline. > > The redirect design structurally avoids the Node.js endpoint-session > bug. Outside the generated Thrift code there is exactly one HashMap > -- the device-to-endpoint hint cache (redirect.rs:68). There is no > endPointToSession map, and Connection::open has exactly two > non-test call sites (session.rs:240, session.rs:319), both of which > only iterate config.endpoints. The client never opens a connection > because of a redirect. Combined with move semantics in the pool > (pop_front at pool.rs:223, remove(pos) at :288, push_back only > after the guard yields ownership at :429-431), a session cannot > simultaneously be in an endpoint map and in the general idle queue. > The comment at pool.rs:270-272 states the trade-off honestly. This > is the single most important thing this client does differently. > > No lock is ever held across blocking I/O. I checked every non-test > lock() in pool.rs: new() :151 (open is at :150, outside), acquire() > :208 (drops at :215 before closing swept sessions, :226/:234 before > hand_out), acquire_for_device() :275 (drops at :289), close() :309 > (scope ends :312, blocking closes at :314-316), release() :365 > (drops at :374), open_session() :324, hand_out() :334. The one > Session destructor that runs under the lock (:228) is a no-op > because connection is already None. sweep_idle being #[must_use] > and returning expired sessions for the caller to close outside the > lock (:173-176) is a deliberate, correct design. > > The borrow checker is doing real work. PooledSession<'a> borrows > the pool, so a guard cannot outlive it; SessionDataSet<'a> borrows > &mut Session, so a result set cannot outlive its session or be used > concurrently with another RPC; every RPC takes &mut self, which > makes interleaved Thrift sequence numbers on one channel impossible > without a runtime lock. Zero unsafe in the crate. > > > Checked, nothing found > ---------------------- > > Recording these because a clean result is information too: async > and cancellation issues (the crate is fully synchronous -- zero > async fn / .await / tokio across src, tests and examples, so the > Node.js-shaped hypotheses are vacuous); RefCell misuse (the two > occurrences are unused generator imports); waiters removed on > timeout but still notified (no explicit waiter list exists); two > callers receiving the same idle session (stress-verified, 16 > threads x 300 acquire/release, zero failures); stale-index > remove/pop after dropping the lock; live accounting drift and usize > underflow (three increments and seven decrements reconcile); mutex > poisoning cascades (no reachable first panic found, though the > chain is fragile); double closeSession / closeOperation; dangling > PooledSession after the pool is dropped; oversized allocation from > a bogus peer; min_size / max_size boundaries; spurious condvar > wakeups; Thrift sequence-number interleaving; Endpoint::parse edge > cases; and redirect interaction with reconnect. > > > Happy to go deeper on any of these, or to re-check a scenario if > you think one is a misread. > > Best regards, > Zihan Dai > GitHub: PDGGK > > On Mon, Jul 20, 2026 12:40 PM, Haonan Hou <[email protected]> wrote: > > > Hi Xuan, > > > > The vote passed and the repo has been created. > > https://github.com/apache/iotdb-client-rust > > > > Thanks, > > Haonan > > > > On 2026/07/17 03:57:02 王旋 wrote: > > > Hi Yuan, > > > > > > Thank you for the suggestions. Status on both: > > > > > > 1. Compatibility & release policy -- now landed in the repository as > > > COMPATIBILITY.md <http://compatibility.md/> [1]: > > > > > > - IoTDB server compatibility matrix: 2.0.1-2.0.8, 2.0.10 and master are > > > listed; 2.0.6 and 2.0.10 are tested (CI live suite / full benchmark + > > > data-correctness verification respectively), the other versions are > > > explicitly marked as untested. > > > - A per-release protocol toolchain table recording, for each crate > > > version, the exact IDL source (currently apache/iotdb master, > > > iotdb-protocol/ @ 2fedd8a395) and the Thrift compiler version > (0.23.0, > > > as pinned by the IoTDB pom), plus the thrift crate version. > > > - A SemVer policy (Cargo semantics, 0.x pre-1.0 rules), a deprecation > > > policy (#[deprecated] for at least one minor release before removal), > > > and a release checklist that updates the toolchain table on every > > > release. > > > > > > CI now runs the live integration suite and examples against both the > > > oldest tested and the latest stable IoTDB (2.0.6 and 2.0.10) as a build > > > matrix [2]. > > > > > > 2. crates.io publishing permissions and a community-managed release > > > process -- fully agreed. Since this depends on the Apache repository > and > > > community processes existing first, I will file it as a roadmap issue > in > > > apache/iotdb-client-rust once the repository is created, covering > > > community-owned crates.io ownership, reproducible release steps, and > > > review by multiple maintainers. > > > > > > [1] > > > > > > https://github.com/CritasWang/iotdb-client-rust/blob/main/COMPATIBILITY.md > > > [2] > > > > > > https://github.com/CritasWang/iotdb-client-rust/blob/main/.github/workflows/ci.yml > > > > > > Best regards, > > > Xuan Wang > > > > > > zerolbsony <[email protected]> 于2026年7月16日周四 20:58写道: > > > > > > > Hi, Xuan Wang > > > > It’s hard to believe that Rust has overtaken Go and become so > popular. > > > > The top 10 programming languages in the July 2026 TIOBE rankings are > as > > > > follows: > > > > Python – 18.94% > > > > C – 10.86% > > > > C++ – 9.12% > > > > Java – 8.03% > > > > C# – 4.49% > > > > JavaScript – 2.72% > > > > Visual Basic – 2.48% > > > > SQL – 1.71% > > > > R – 1.69% > > > > Rust – 1.34% > > > > > > > > > > > > I’ve noticed that starting from 2024, the Linux kernel has allowed > > certain > > > > modules, primarily device drivers, to be written in Rust. > > > > > > > > > > > > Rust can be used to write Linux kernel modules, which means you can > run > > > > Rust code directly at the kernel level. This is extremely useful for > > > > scenarios that demand high performance and low-level hardware > > operations. > > > > Have you considered invoking certain IoTDB functionalities directly > > via the > > > > Linux kernel to boost performance? > > > > Best regards, > > > > Bo Li > > > > > > > > > > > > >Hi, > > > > > > > > > >Thank you very much for the supportive and detailed reply. The > > suggested > > > > >route (code donation to the existing TLP + ASF IP clearance, rather > > than > > > > >incubation) makes sense to me, and I'm glad full Java-parity is not > a > > > > >prerequisite — an incremental, roadmap-tracked approach matches how > > the > > > > >client was built. > > > > > > > > > >Below are the five items requested before a vote. > > > > > > > > > >## 1. Supported IoTDB version matrix > > > > > > > > > >| IoTDB version | Status | Evidence | > > > > >|---|---|---| > > > > >| 2.0.6 | Supported | CI integration job runs the full live test > > suite + 3 > > > > >examples against apache/iotdb:2.0.6-standalone on every push | > > > > >| 2.0.10 | Supported | Full benchmark + data-correctness > verification > > (all > > > > >10 data types, tree & table, nulls) against a 2.0.10 standalone > > > > deployment | > > > > >| 2.0.x (general) | Expected to work | Thrift IDL is synced from > > > > >iotdb-protocol/ master; protocol version V3; no version-specific > > branching > > > > >in the client | > > > > >| 1.x | Not targeted | Table model and several data types > > > > >(TIMESTAMP/DATE/BLOB/STRING) assume 2.x; no plan to support 1.x | > > > > > > > > > >TLS was verified end-to-end against 2.0.6 with enable_thrift_ssl > > (TLSv1.3, > > > > >full certificate verification). RPC compression (compact protocol) > > > > verified > > > > >against a server with dn_rpc_thrift_compression_enable=true. > > > > > > > > > >## 2. Contributors and provenance > > > > > > > > > >The entire git history has a single author — myself (Wang Xuan / > > > > CritasWang > > > > ><[email protected]>). I am an existing Apache IoTDB committer with > > an > > > > ICLA > > > > >already on file, and I hold full ownership of this work, so the > > provenance > > > > >chain is straightforward. Two categories of code in the repo: > > > > > > > > > >- Hand-written code: 100% original work, written for this project, > > > > >Apache-2.0 headers on every file. > > > > >- Generated code (src/protocol/): produced by the Apache Thrift > > compiler > > > > >from the Apache IoTDB project's own IDL files (iotdb-protocol/, > > ALv2); the > > > > >generation pipeline (tools/generate-thrift.sh) is documented and > > > > >reproducible. > > > > > > > > > >No code was copied from other client SDKs; the Java/C#/Node.js > clients > > > > were > > > > >used as behavioral references only (wire-protocol semantics), which > is > > > > also > > > > >how the cross-client protocol issues (e.g. the Node.js DATE encoding > > bug) > > > > >were found. > > > > > > > > > >## 3. Dependency and license inventory > > > > > > > > > >Runtime dependencies (4 + 1 optional): > > > > > > > > > >| Crate | Version | License | Purpose | > > > > >|---|---|---|---| > > > > >| thrift | 0.23 | Apache-2.0 | RPC (matches the compiler version > > pinned by > > > > >the IoTDB pom) | > > > > >| byteorder | 1.5 | Unlicense OR MIT | Big-endian encoding | > > > > >| chrono | 0.4 | MIT OR Apache-2.0 | DATE handling | > > > > >| log | 0.4 | MIT OR Apache-2.0 | Logging facade | > > > > >| native-tls (optional, feature "tls") | 0.2 | MIT OR Apache-2.0 | > > TLS | > > > > > > > > > >Dev-dependency: env_logger (MIT OR Apache-2.0). Full transitive > > closure: > > > > 74 > > > > >packages, machine-checked — every one is MIT / Apache-2.0 / > Unlicense > > / > > > > BSD > > > > >/ Zlib dual-or-multi licensed, i.e. 100% ASF Category-A. Zero > > > > Category-B/X, > > > > >zero unknown-license packages. The full inventory (cargo metadata) > > can be > > > > >attached to the IP-clearance checklist. > > > > > > > > > >## 4. Proposed initial maintainers and maintenance plan > > > > > > > > > >- Initial maintainer: myself (CritasWang — existing IoTDB > committer). > > I > > > > >commit to maintaining the client — issue triage, protocol-sync with > > > > >iotdb-protocol changes, and regular releases through the normal > IoTDB > > > > >community process — and to growing co-maintainers from the > community; > > > > >review bandwidth from other committers/PMC members on client-facing > > > > changes > > > > >would be welcome. > > > > >- The repo already carries the practices expected for handover: CI > > (fmt / > > > > >clippy -D warnings / license check / unit tests / live integration > > against > > > > >a service container), bilingual READMEs, runnable examples, a > > > > >statistics-aligned benchmark, and a documented codegen pipeline. > > Nothing > > > > >depends on my personal infrastructure. > > > > > > > > > >## 5. API support matrix (vs Java client) and near-term roadmap > > > > > > > > > >Supported today (feature → Java-client equivalent): > > > > > > > > > >| Area | Rust today | Java client | > > > > >|---|---|---| > > > > >| Session lifecycle, multi-node URLs, failover, auto-reconnect | ✅ | > > ✅ | > > > > >| insertTablet / insertTablets | ✅ | ✅ | > > > > >| insertRecord(s) / insertRecordsOfOneDevice (+ aligned variants) | > ✅ > > | ✅ > > > > | > > > > >| Table model (TableSession/Pool, TAG/FIELD/ATTRIBUTE) | ✅ | ✅ | > > > > >| SessionPool / TableSessionPool (idle eviction, redirection cache) > | > > ✅ | > > > > ✅ > > > > >| > > > > >| Query + paging iteration (TsBlock, fetchResultsV2) | ✅ | ✅ | > > > > >| All data types incl. TIMESTAMP/DATE/BLOB/STRING | ✅ | ✅ | > > > > >| TLS (incl. client identity) / RPC compression | ✅ | ✅ | > > > > >| deleteData / deleteTimeseries / DDL helpers | SQL only | ✅ > dedicated > > > > APIs > > > > >| > > > > >| Schema templates | ❌ | ✅ | > > > > >| executeRawDataQuery / executeLastDataQuery / aggregation APIs | ❌ > | > > ✅ | > > > > >| Per-type encoding tuning (RPC compression V2) | ❌ | ✅ | > > > > > > > > > >Near-term roadmap (order negotiable with the community): > > > > > > > > > >1. async (tokio) API layer — most-requested Rust-ecosystem feature, > > kept > > > > >out of v0.1 to keep the sync core small > > > > >2. crates.io releases under the Apache project > > > > > > > > > >I'll keep this thread open for feedback on maintenance ownership and > > API > > > > >scope as suggested. Since I'm an existing committer with an ICLA on > > file > > > > >and the sole author, the IP-clearance paperwork should be light — > I'm > > > > happy > > > > >to prepare whatever the PMC deems necessary (software grant if > > required > > > > for > > > > >the pre-existing external history, plus the dependency inventory > > above) > > > > >whenever the community feels ready to move to a VOTE. > > > > > > > > > >Best regards, > > > > >Xuan Wang > > > > > > > > > >Haonan Hou <[email protected]> 于2026年7月15日周三 15:06写道: > > > > > > > > > >> Hi Xuan, > > > > >> > > > > >> Thank you for sharing this work. The implementation already > covers a > > > > >> substantial part of the client functionality, and the test > coverage, > > > > live > > > > >> integration testing, CI setup, and benchmark results are all > > > > encouraging. > > > > >> > > > > >> The protocol issues and performance improvements discovered while > > > > >> validating > > > > >> against the other clients are also a good indication that this > work > > can > > > > >> benefit the broader IoTDB client ecosystem. > > > > >> > > > > >> As a member of the IoTDB PMC, I support bringing a well-maintained > > Rust > > > > >> client > > > > >> into the Apache IoTDB project. Based on what you have presented, I > > > > believe > > > > >> this > > > > >> repository is a strong candidate for donation and for becoming an > > > > official > > > > >> IoTDB client. > > > > >> > > > > >> Since IoTDB is already an Apache top-level project, the > appropriate > > > > route > > > > >> should be a code donation to the existing project, followed by the > > ASF > > > > IP > > > > >> clearance process, rather than incubation as a separate project. > > > > >> > > > > >> I suggest the following next steps: > > > > >> > > > > >> 1. Continue this DISCUSS thread to collect feedback from the > > community, > > > > >> especially regarding maintenance ownership and the initial API > > scope. > > > > >> 2. Prepare a concise compatibility and API-support matrix against > > the > > > > Java > > > > >> client, together with a near-term roadmap. > > > > >> 3. If there are no major concerns, start a formal PMC VOTE to > > accept the > > > > >> code > > > > >> donation and create the apache/iotdb-client-rust repository. > > > > >> 4. Complete the software grant, code provenance, > dependency/license > > > > review, > > > > >> and ASF IP clearance. > > > > >> 5. Move the code to the Apache repository and continue development > > and > > > > >> releases through the normal IoTDB community process. > > > > >> > > > > >> I do not think complete Java client API parity needs to be a > > > > prerequisite > > > > >> for > > > > >> the donation. The current functionality appears sufficient to > > establish > > > > a > > > > >> useful official client. APIs such as schema-template operations > and > > > > >> additional > > > > >> session-level query methods can be added incrementally and tracked > > > > through > > > > >> a > > > > >> public roadmap. > > > > >> > > > > >> Before moving to a vote, it would be helpful to provide: > > > > >> > > > > >> - The supported IoTDB version matrix. > > > > >> - A list of contributors and confirmation of the code’s > > > > >> ownership/provenance. > > > > >> - A dependency and license inventory. > > > > >> - The proposed initial maintainers and long-term maintenance plan. > > > > >> - An API support matrix and near-term roadmap. > > > > >> > > > > >> Overall, I am supportive of this proposal and am willing to help > > move > > > > the > > > > >> discussion and donation process forward. > > > > >> > > > > >> Best regards, > > > > >> Haonan Hou > > > > >> > > > > >> On 2026/07/15 03:50:24 王旋 wrote: > > > > >> > Hi all, > > > > >> > > > > > >> > I'd like to share a Rust client SDK for Apache IoTDB that I've > > been > > > > >> > developing, and ask for the community's feedback on whether > there > > is > > > > >> > interest in bringing it into the Apache IoTDB ecosystem. > > > > >> > > > > > >> > Repository: https://github.com/CritasWang/iotdb-client-rust > > > > >> > > > > > >> > Scope & features > > > > >> > > > > > >> > Tree model (Session) and table model (TableSession / SQL > dialect), > > > > >> > mirroring the API shape of the Java / C# / Node.js clients > > > > >> > Write APIs: insertTablet / insertTablets (multi-tablet > batching), > > > > >> > insertRecord(s), insertRecordsOfOneDevice, plus aligned variants > > > > >> > SessionPool and TableSessionPool: RAII checkout, lazy growth, > idle > > > > >> > eviction, dead-connection eviction, write-redirection cache > > (status > > > > 400 > > > > >> > redirect hints), automatic reconnect with endpoint failover > > > > >> > Full data-type coverage including TIMESTAMP / DATE / BLOB / > > STRING; > > > > >> TsBlock > > > > >> > decoding with logical-type re-tagging (the TsBlock header > carries > > > > >> physical > > > > >> > types — DATE arrives as INT32, BLOB as TEXT) > > > > >> > TLS (feature-gated, incl. client identity) — verified end-to-end > > > > against > > > > >> a > > > > >> > real IoTDB with enable_thrift_ssl; RPC compression (compact > > protocol) > > > > >> > Thrift codegen pipeline sources the IDL from iotdb-protocol/ and > > uses > > > > the > > > > >> > thrift compiler fetched by the IoTDB Maven build, so the stubs > > stay in > > > > >> > lockstep with the server > > > > >> > > > > > >> > Quality > > > > >> > > > > > >> > 113 unit tests (124 with TLS) + live integration tests that skip > > > > >> gracefully > > > > >> > without a server; CI on GitHub Actions runs fmt / clippy -D > > warnings / > > > > >> > license checks plus an integration job against an IoTDB 2.0.6 > > service > > > > >> > container — green > > > > >> > Statistics-aligned benchmark (measurement semantics mirror > > > > iot-benchmark: > > > > >> > prep inside the timed span, failures excluded from latency, > > > > >> Result/Latency > > > > >> > Matrix output). On a 16-core server writing 2B points per run > > (table > > > > >> model, > > > > >> > 100 devices x 20 DOUBLE sensors, 1000-row tablets, 20 sessions), > > the > > > > Rust > > > > >> > client sustains ~46-47M points/s — statistically tied with > > > > >> iot-benchmark's > > > > >> > Java Session path on the same box, where the server, not the > > client, > > > > is > > > > >> the > > > > >> > ceiling > > > > >> > Every file carries the ASF Apache 2.0 header > > > > >> > > > > > >> > Side effects the community already received > > > > >> > > > > > >> > Cross-validating the wire protocol against the Java / C# / > Node.js > > > > >> > implementations surfaced two upstream issues in the Node.js > > client, > > > > both > > > > >> > now addressed: the DATE wire-encoding fix > > > > (apache/iotdb-client-nodejs#14, > > > > >> > PR #15, merged) and a write-path serialization optimization > (+52% > > > > >> > throughput, PR #16, under review). > > > > >> > > > > > >> > Questions for the community > > > > >> > > > > > >> > Is there interest in an official Rust client under the Apache > > IoTDB > > > > >> > umbrella (e.g., apache/iotdb-client-rust), following the path of > > the > > > > >> > Node.js / C# clients? > > > > >> > If so, what would the preferred route be — code donation via the > > > > >> incubator > > > > >> > process for client SDKs, or starting a repo under the existing > > project > > > > >> and > > > > >> > iterating there? > > > > >> > Any API-surface expectations from the PMC side before such a > move > > > > (e.g., > > > > >> > schema-template APIs, session-level query APIs like > > > > executeRawDataQuery)? > > > > >> > I'm happy to keep maintaining it either way, and to align the > > roadmap > > > > >> with > > > > >> > the community's priorities. > > > > >> > > > > > >> > 大家好 > > > > >> > > > > > >> > 我想向社区分享一个我一直在开发的 Apache IoTDB Rust 客户端 SDK,并征求大家的意见:社区是否有兴趣将它纳入 > > Apache > > > > >> > IoTDB 生态。 > > > > >> > > > > > >> > 仓库地址:https://github.com/CritasWang/iotdb-client-rust > > > > >> > > > > > >> > 范围与功能 > > > > >> > > > > > >> > 树模型(Session)与表模型(TableSession / SQL 方言),API 形态与 Java / C# / > > Node.js > > > > 客户端对齐 > > > > >> > 写入 API:insertTablet / insertTablets(多 tablet > > > > >> > 批量)、insertRecord(s)、insertRecordsOfOneDevice,以及 aligned 变体 > > > > >> > SessionPool 与 TableSessionPool:RAII > > 借还、惰性增长、空闲回收、死连接剔除、写重定向缓存(status > > > > 400 > > > > >> > 重定向提示)、带端点故障转移的自动重连 > > > > >> > 完整数据类型覆盖,包括 TIMESTAMP / DATE / BLOB / STRING;TsBlock > > > > 解码带逻辑类型重标记(TsBlock > > > > >> > 头携带的是物理类型——DATE 以 INT32 到达、BLOB 以 TEXT 到达) > > > > >> > TLS(feature 门控,含客户端证书)——已对开启 enable_thrift_ssl 的真实 IoTDB > > 完成端到端验证;RPC > > > > >> > 压缩(compact 协议) > > > > >> > Thrift 代码生成流水线:IDL 取自 iotdb-protocol/,编译器使用 IoTDB Maven > 构建拉取的版本,确保 > > > > stub > > > > >> > 与服务端严格同步 > > > > >> > > > > > >> > 质量 > > > > >> > > > > > >> > 113 个单元测试(含 TLS 为 124 个)+ 无服务器时优雅跳过的 live 集成测试;GitHub Actions CI > > 运行 > > > > fmt / > > > > >> > clippy -D warnings / license 检查,另有针对 IoTDB 2.0.6 service > > container 的集成 > > > > >> > job——全绿 > > > > >> > 统计口径对齐的基准测试(测量语义对齐 iot-benchmark:批准备计入计时段、失败不计入延迟、Result/Latency > > > > Matrix > > > > >> > 输出)。在 16 核服务器上每轮写入 20 亿点(表模型、100 设备 x 20 个 DOUBLE 测点、1000 > > 行/tablet、20 > > > > >> > 会话),Rust 客户端持续吞吐约 4600-4700 万点/秒——与同机 iot-benchmark 的 Java > Session > > > > >> > 路径统计学持平,此时瓶颈在服务端而非客户端 > > > > >> > 所有文件均带 ASF Apache 2.0 头 > > > > >> > > > > > >> > 社区已经收到的副产品 > > > > >> > > > > > >> > 在与 Java / C# / Node.js 实现交叉验证线上协议的过程中,发现了 Node.js > > > > 客户端的两个上游问题,目前均已处理:DATE > > > > >> > 线上编码修复(apache/iotdb-client-nodejs#14,PR #15,已合并)以及写路径序列化优化(吞吐 > > +52%,PR > > > > >> > #16,评审中)。 > > > > >> > > > > > >> > 想请教社区的问题 > > > > >> > > > > > >> > 社区是否有兴趣在 Apache IoTDB 旗下提供官方 Rust 客户端(例如 > > apache/iotdb-client-rust),沿用 > > > > >> > Node.js / C# 客户端的路径? > > > > >> > 如果有,倾向的路线是什么——通过客户端 SDK 的代码捐赠流程,还是先在现有项目下建仓库迭代? > > > > >> > 在此之前,PMC 对 API 面是否有预期要求(例如 schema 模板 API、executeRawDataQuery > > 等会话级查询 > > > > API)? > > > > >> > 无论结果如何,我都会继续维护它,并愿意将路线图与社区的优先级对齐。 > > > > >> > > > > > >> > Best regards, > > > > >> > Xuan Wang > > > > >> > > > > > >> > > > > > > > > > >
