quantranhong1999 opened a new pull request, #3208:
URL: https://github.com/apache/james-project/pull/3208

   cc @felixauringer @chibenwa 
   
   ### Problem
   A throttled full re-indexing on the Postgres app 
(`/mailboxes?task=reIndex&messagesPerSecond=N`)
   always fails after exactly `jooq.reactive.timeout` (10s by default), having 
indexed 9 x N messages:
   
   ```
   PostgresExecutor - Time out executing Postgres query. May need to check 
either jOOQ reactive issue or Postgres DB performance.
   java.util.concurrent.TimeoutException: Did not observe any item or terminal 
signal within 10000ms in 'flatMapMany'
   ```
   
   Reported on server-dev and in linagora/tmail-backend#1599, where the timeout 
is followed by
   further Postgres failures until the server "sometimes recovers".
   
   ### Root cause
   Not a hanging query. The full re-indexing lists all mailboxes with a mailbox 
concurrency of 1 and
   throttles the messages of each mailbox, so the mailbox listing query stays 
open while the first
   mailbox is slowly re-indexed. `PostgresExecutor` guarded streamed results 
with `Flux.timeout`,
   whose timer restarts on every emitted row regardless of whether the consumer 
requested more rows.
   A consumer slower than the timeout is therefore failed while Postgres is 
only waiting for demand.
   
   A second issue was found while reproducing: cancelling the pipeline on 
timeout does not stop the
   query on the Postgres side. The busy connection goes back to the pool, the 
next query on it queues
   behind the still running statement and times out too, which is the cascade 
seen in the field.
   
   ### Fixes
   - **Demand aware timeout**: new `TimeoutOnPendingDemand` operator, used for 
streamed Flux results.
     The timer only runs while the downstream has pending demand and restarts 
on each delivered row.
     A genuine stall while rows are awaited still raises `TimeoutException`. 
Mono results are unchanged.
   - **Cancel on timeout**: on `TimeoutException`, `PostgresExecutor` sends a 
r2dbc-postgresql
     `cancelRequest()` before releasing the connection, so it is reusable 
within milliseconds.
   - **Library upgrade**: jOOQ 3.20.5 -> 3.20.20 and r2dbc-postgresql 1.0.7 -> 
1.1.3, which ship the
     fixes for cancelled queries poisoning pooled connections (jOOQ#20112, 
r2dbc-postgresql#661).
   
   ### Tests (TDD, reproducing tests committed before each fix)
   - `PostgresReIndexingIntegrationTest`: the reported webadmin scenario on a 
Guice Postgres server.
   - `PostgresExecutorTimeoutTest`: slow consumer must not time out, a stalled 
query must, and a
     single-connection pool is usable right after a timeout (60s `pg_sleep`).
   - `TimeoutOnPendingDemandTest`: 10 virtual-time cases for the operator.
   
   Validated on the Postgres backend, mailbox, blob, data, JMAP data, 
event-bus, task and vault
   module suites, and the full Postgres webadmin integration module.
   
   ### Note for reviewers
   As discussed on the list, paginating the mailbox listing (per-page queries) 
would additionally avoid
   holding a pooled connection during slow consumption, and would preserve the 
previous accidental
   "deadlock breaker" behaviour of the timeout under pool exhaustion. This PR 
keeps streaming and can
   be combined with that follow-up.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to