wangmingzhou1986 opened a new issue, #67898:
URL: https://github.com/apache/doris/issues/67898

   ### Search before asking
   
   - [x] I had searched in the issues and found no similar issues.
   
   ### Version
   
   `doris-4.1.3-rc02(AVX2) RELEASE`
   features: 
`-TDE,-VARIANT_NESTED_GROUP,-HDFS_STORAGE_VAULT,+UI,+AZURE_BLOB,+AZURE_STORAGE_VAULT,-HIVE_UDF,+BE_JAVA_EXTENSIONS`
   build: `git://vm-122@7126cf65d96ebc43fce0906f51e92c1a2ccf24a6`
   
   ### What's Wrong?
   
   Every query against an **external catalog** (Paimon on S3/MinIO, via `CREATE 
CATALOG ... 'type'='paimon'`)
   leaks roughly **10 threads per BE**, named `rs_normal`. The threads are 
never reclaimed.
   Thread count therefore grows linearly with the number of external-table 
queries,
   and because jemalloc's per-thread `tcache` scales with thread count,
   BE `rss` grows with it until it hits `mem_limit` and **every** query fails 
with
   `MEM_LIMIT_EXCEEDED` — while `SHOW BACKENDS` still reports `Alive: true`.
   
   Setting `doris_max_remote_scanner_thread_pool_thread_num` to a finite value
   (we tried `256`, down from the default `-1`) and restarting BE **does not 
bound this pool**:
   the count climbed past 3,192 and kept going.
   
   ### Evidence
   
   #### 1. Only `rs_normal` grows
   
   Thread-name histogram from `/proc/<be_pid>/task/*/comm`:
   
   | thread name | be-0 (uptime 17.7 h) | be2-0 (uptime 22 h) |
   |---|---|---|
   | **`rs_normal`** | **24,377** | **35,332** |
   | `brpc_arrow_flig` | 512 | 512 |
   | `doris_be` | 352 | 344 |
   | `brpc_light` / `brpc_heavy` / `EvHttpServer` | 128 / 128 / 128 | 128 / 128 
/ 128 |
   | `SendBatchThreadP` / `DownloadThreadP` | 64 / 64 | 64 / 64 |
   | `ls_normal` | 48 | 48 |
   | `p_normal_blocki` / `TabletPublishTx` / `SegmentPrefetch` | 32 / 32 / 32 | 
32 / 32 / 32 |
   | everything else, summed | ~1,779 | ~1,771 |
   
   Every other pool is identical between the two BEs and stays flat. Only 
`rs_normal` diverges,
   and it diverges in proportion to uptime (i.e. to accumulated query count).
   
   #### 2. Controlled experiment — growth is caused by external-catalog 
queries, and is not reclaimed
   
   Both BEs observed simultaneously; `rs_normal` counted before and after each 
phase:
   
   | phase | be-0 Δ | be2-0 Δ |
   |---|---|---|
   | idle 60 s (zero queries) | **-1** | **0** |
   | 20 × `SELECT COUNT(*) FROM <paimon_catalog>.<db>.<tbl>` | **+203** | 
**+201** |
   | idle 120 s | **-4** | **-1** |
   
   ≈ **10 threads per query per BE**. Queries fan out to all BEs, so both grow 
together.
   Idle neither grows nor reclaims.
   
   #### 3. The config knob does not bound it
   
   From `/api/show_config` (all three are `mutable=false`):
   
   ```
   doris_remote_scanner_thread_pool_thread_num       = 48
   doris_max_remote_scanner_thread_pool_thread_num   = -1     <- default in 
this build
   doris_remote_scanner_thread_pool_queue_size       = 102400
   ```
   
   Note the docs state the default for 
`doris_max_remote_scanner_thread_pool_thread_num`
   is **512**, but this build reports **-1**.
   
   We set it explicitly to `256` in `be.conf` and restarted the BE. Verified it 
took effect:
   
   ```
   doris_max_remote_scanner_thread_pool_thread_num = 256
   ```
   
   `rs_normal` nevertheless climbed past **3,192** and kept growing at the same 
rate.
   So whatever creates `rs_normal` threads is not governed by this parameter.
   
   Ruled out: **workload groups are not the cause.** We have exactly **one** 
workload group,
   and `enable_workload_group_for_scan = false`.
   
   #### 4. Memory consequence
   
   jemalloc stats at `rss` 6.81 GB (be2-0, 35 k threads):
   
   | metric | be-0 | be2-0 |
   |---|---|---|
   | `jemalloc_allocated_bytes` | 2.49 GB | 4.02 GB |
   | **`jemalloc_tcache_bytes`** | 1.09 GB | **1.51 GB** |
   | **`jemalloc_metadata_bytes`** | 0.77 GB | **1.08 GB** |
   | `jemalloc_retained_bytes` | 1.12 GB | 1.52 GB |
   
   35 k threads × ~40 KB tcache ≈ 1.5 GB, matching the measured `tcache` almost 
exactly.
   `tcache + metadata` = 2.59 GB = **38 %** of `rss`.
   
   Eventually:
   
   ```
   [MEM_LIMIT_EXCEEDED] ... process memory used 6.62 GB(= 6.62 GB[vm/rss]),
   limit 7.00 GB, soft limit 6.30 GB, sys available memory 398 MB
   ```
   
   At that point *every* query through the BE fails, including `SELECT ... 
LIMIT 1`,
   yet `SHOW BACKENDS` still shows `Alive: true` — so health checks based on 
liveness
   do not detect it.
   
   Restarting BE clears it completely, confirming the threads (not data) hold 
the memory:
   
   | | before restart | after restart |
   |---|---|---|
   | be2-0 rss / threads | 6,805 MB / 37,103 | 971 MB / 1,778 |
   | be-0 rss / threads | 4,514 MB / 26,579 | 1,361 MB / 1,933 |
   
   #### 5. Environment note (may or may not be relevant)
   
   BE runs in a container. `nproc` inside the container reports **16** (the 
host's core count),
   while the cgroup CPU limit is **2** (be2-0) / **6** (be-0).
   
   ```
   JEMALLOC_CONF=percpu_arena:percpu,background_thread:true,metadata_thp:auto,
   muzzy_decay_ms:5000,dirty_decay_ms:5000,oversize_threshold:0,prof:true,
   prof_active:false,lg_prof_interval:-1,lg_extent_max_active_fit:8
   ```
   
   `dirty_decay_ms`/`muzzy_decay_ms` are already aggressive (5 s), so this is 
not jemalloc
   withholding freed pages — the memory is genuinely attached to live threads.
   
   ### What You Expected?
   
   `rs_normal` threads are returned to (or bounded by) a pool after the scan 
finishes,
   so that BE thread count and `rss` stay flat under a steady stream of 
external-catalog
   queries, and `doris_max_remote_scanner_thread_pool_thread_num` actually caps 
the pool.
   
   ### How to Reproduce?
   
   1. Create an external catalog (we used Paimon on S3; a Hive/Iceberg catalog 
on object
      storage will likely do as well).
   2. Record the baseline: `for t in /proc/<be_pid>/task/*/comm; do cat $t; 
done | grep -c '^rs_normal'`
   3. Run N (e.g. 20) trivial queries against a table in that catalog,
      e.g. `SELECT COUNT(*) FROM <catalog>.<db>.<tbl>;`
   4. Re-count `rs_normal`. Expect ≈ `10 × N` more threads per BE.
   5. Wait a few minutes with no queries and count again — the threads are not 
reclaimed.
   
   ### Anything Else?
   
   - Happy to provide the full thread-name histograms, `/api/show_config` dumps,
     `/metrics` jemalloc series, or a time series of `rss` / thread count
     (we sample both every 5 minutes).
   - If `rs_normal` is expected to be bounded by a different config than
     `doris_max_remote_scanner_thread_pool_thread_num`, please point us at it 
and
     we'll re-test and report back.
   - The discrepancy between the documented default (`512`) and this build's 
default (`-1`)
     may itself be worth a look.
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to