Rangsh commented on issue #12058:
URL: https://github.com/apache/seatunnel/issues/12058#issuecomment-5674902093

   @SEZ9 @DanielLeens @nzw921rx following up with the wait-site experiment on 
**unchanged** `upstream/dev` (flush-path remains ruled out; #12081 stays 
Related-only).
   
   Artifacts: https://gist.github.com/Rangsh/d7129f0f422dad61da1130e53ebfb0ba 
(`WAIT-SITE-REPORT.md` + flames / primary JMH JSON / EXTRA log / park-summary / 
0 ns ThreadPark `.jfc`). The raw `.jfr` is binary so it is not in the gist; the 
parsed park numbers are in `meta/park-summary.json` and the report — happy to 
share the `.jfr` another way if useful.
   
   ### Controlled question
   
   Is `InvocationFuture.get` / `LockSupport.park` a **causal wait** for 
`checkpointOverviewIncrementalUpdate$`, or a **passive symptom**?
   
   Treated as a wait site until the data says otherwise. Write-through + caller 
completion wait preserved. `fs.defaultFS: file:///` (filesystem-specific).
   
   ### Environment
   
   | Item | Value |
   | --- | --- |
   | Commit | `38a37104e` (`upstream/dev`) |
   | Machine | Apple M1 / macOS 14.7.x |
   | JDK | Corretto **11.0.26** |
   | Method | `CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate$` 
|
   | JMH | existing defaults (`SingleShotTime`, 3 forks, 3/5 
warmup/measurement, `@OperationsPerInvocation(100)`) |
   | Store | write-through MapStore, `file:///` |
   
   ### Primary JMH (raw per-iteration)
   
   **164.511 ± 16.799 us/op**, CV **9.55%**, min–max **123.7–197.1** (quiet 
day; not the historical ~32% / ~393 us spike day).
   
   Raw forks (us/op):
   
   - fork 0: 160.59, 172.19, **197.08**, 159.94, 154.82
   - fork 1: 173.84, 164.72, 161.38, 123.74, 163.33
   - fork 2: 183.21, 166.95, 168.55, 160.38, 156.95
   
   ### Wake-up / notification + scheduler timing
   
   1. **Wall/CPU flames** (diagnostic `-f 1`): timed path still  
      `updateOverview` → `MapProxyImpl.compute` → 
`AbstractInvocationFuture.get` → `LockSupport.park`, with write-through 
`RequestFuture.get` / `CountDownLatch.await` on the partition-operation thread. 
 
      **Wake path observed:** `LockSupport.unpark` → 
`AbstractInvocationFuture.unblockAll` → invocation `complete` / `sendResponse` 
(not an anonymous park).
   
   2. **Low-threshold JFR `jdk.ThreadPark` (0 ns)** — default `profile` JFR 
hides µs parks; with threshold 0 ns on the same method:
   
   | Filter | n | mean | CV | p50 | p95 | max |
   | --- | ---: | ---: | ---: | ---: | ---: | ---: |
   | jmh-worker × `compute` × `AbstractInvocationFuture` | 1200 | **57.7 µs** | 
133.9% | 55.5 | 128.5 | **1851 µs** |
   | partition-op × `RequestFuture` / latch await | 1917 | **33.9 µs** | 309.6% 
| 20.5 | 83.2 | **2955 µs** |
   
   Typical InvocationFuture park (~58 µs) is the same order as the primary op 
(~165 µs): the wait is a **large share of the timed path**, not a zero-duration 
label. Heavy tails exist (ms-scale parks), which remain a **plausible** driver 
for high-CV days under `@OperationsPerInvocation(100)` (one 1.85 ms park ≈ 
+18.5 µs/op on that sample), but this quiet run does **not** by itself prove 
the historical 393 µs outliers.
   
   GC young pauses are present; still no tight iteration↔pause proof on this 
quiet day.
   
   ### Verdict (no production conclusion)
   
   - **Wait site confirmed**; sync/`hsync` still not the CPU-hot driver.
   - **Lean causal for steady-state timed-path cost** (park duration ≈ main 
measured work).
   - **Spike-day causality still open** until a high-CV run is park-sum aligned 
(or a labelled `OperationsPerInvocation(1)` diagnostic).
   - **No production change** proposed. #12081 cannot claim to address #12058 
from this.
   
   Happy to take the noisy-day / per-iteration park-sum tightening if you want 
that next; otherwise this is the wait-site evidence pack for review.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to