wenzhenghu opened a new pull request, #68623:
URL: https://github.com/apache/doris/pull/68623

   ### What problem does this PR solve?
   
   Issue Number: none
   
   Problem Summary:
   
   On many-core hosts, glibc's ptmalloc creates up to `8 × nproc` malloc 
arenas, and every arena is backed by a 64MB-aligned `mmap` segment. Freed 
memory inside an arena is almost never returned to the OS (top-of-heap `trim` 
rarely succeeds once fragmentation sets in), so the FE process RSS retains 
**all historical peak native allocations** — and none of them are visible to 
JVM statistics (heap, non-heap, direct buffers, GC), because native `malloc` 
memory does not belong to any JVM memory pool. The gap keeps growing with the 
number of cores and with heap size, and it is very hard to diagnose precisely 
because every JVM-side metric looks normal.
   
   ### Production case (Doris 3.1.x, 128-core host, `-Xms=-Xmx=200G`)
   
   - FE RSS grew to **451.6 GB** over 33 days while every JVM-visible metric 
stayed "normal":
     - heap committed 200 GB / used 142–154 GB (within `-Xmx`)
     - non-heap 466 MB, direct buffers 1.2 GB (declining), 2,357 threads ≈ 2.4 
GB stacks
   - `/proc/<pid>/smaps` showed **3,227 anonymous 64MB-aligned segments = 410.4 
GB, 90.9% of all anonymous memory** — the signature footprint of glibc arenas
   - Old GC frequency accelerated over two weeks (0/day → 12.5/day → 30/day); 
every GC triggers malloc/free storms on the HotSpot C++ side (G1 RSet / 
PerRegionTable / marking stacks), further fragmenting the arenas
   - The same FE survived an earlier Full-GC incident with a different trigger, 
in which a restart accidentally reset the arenas — without a restart they 
simply accumulated again
   
   This is the same problem Apache Hadoop upstream identified years ago: 
[HADOOP-7154](https://issues.apache.org/jira/browse/HADOOP-7154) landed 
`MALLOC_ARENA_MAX=4` in `hadoop-config.sh`, concluding *"no noticeable downside 
in performance — we've been recommending MALLOC_ARENA_MAX=4"*. HBase handled it 
via [HBASE-5529]. Oracle's JDK team also documents the pattern 
([inside.java](https://inside.java/2023/05/30/usedynamicnumberofcompilerthreads)):
 *"For a machine with a large number of CPUs, the potential maximum number of 
arenas can be very high. Setting MALLOC_ARENA_MAX to a lower value ... can help 
limit fragmentation and prevent an unwanted increase in RSS."*
   
   ### Why 4 arenas is safe for FE
   
   - Java object allocation goes through TLAB, which is JVM-internal and does 
**not** use glibc malloc — the cap only bounds JNI/native-library allocations.
   - Hadoop has shipped `MALLOC_ARENA_MAX=4` as a default for years with no 
throughput regression reported.
   - Users who need more arenas can still override by exporting 
`MALLOC_ARENA_MAX` before running `start_fe.sh` (this PR uses the `${VAR:-4}` 
form to respect pre-set values).
   
   ### What changes are included?
   
   `bin/start_fe.sh`: export `MALLOC_ARENA_MAX=4` (no-op if already set) before 
launching the FE JVM. The change is a pure environment-variable export; no JVM 
flags, no SQL/protocol semantics, and no config keys are touched.
   
   ### Release note
   
   Bound FE native-memory (glibc arena) footprint on many-core hosts via a 
default `MALLOC_ARENA_MAX=4`, overridable by environment.
   
   ### Testing done
   
   - Manual test: no functional code change (environment-variable export in 
`start_fe.sh` before JVM launch). Request reviewers running FE on many-core 
hosts to verify startup and RSS behavior.
   
   ### Check List (For Author)
   
   - Test: Manual test (startup + environ verification); suggest community 
verification on 64+ core hosts
   - Behavior changed: No (environment-variable default only; overridable by 
users)
   - Does this need documentation: No
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to