quantranhong1999 opened a new pull request, #3208:
URL: https://github.com/apache/james-project/pull/3208
cc @felixauringer @chibenwa
### Problem
A throttled full re-indexing on the Postgres app
(`/mailboxes?task=reIndex&messagesPerSecond=N`)
always fails after exactly `jooq.reactive.timeout` (10s by default), having
indexed 9 x N messages:
```
PostgresExecutor - Time out executing Postgres query. May need to check
either jOOQ reactive issue or Postgres DB performance.
java.util.concurrent.TimeoutException: Did not observe any item or terminal
signal within 10000ms in 'flatMapMany'
```
Reported on server-dev and in linagora/tmail-backend#1599, where the timeout
is followed by
further Postgres failures until the server "sometimes recovers".
### Root cause
Not a hanging query. The full re-indexing lists all mailboxes with a mailbox
concurrency of 1 and
throttles the messages of each mailbox, so the mailbox listing query stays
open while the first
mailbox is slowly re-indexed. `PostgresExecutor` guarded streamed results
with `Flux.timeout`,
whose timer restarts on every emitted row regardless of whether the consumer
requested more rows.
A consumer slower than the timeout is therefore failed while Postgres is
only waiting for demand.
A second issue was found while reproducing: cancelling the pipeline on
timeout does not stop the
query on the Postgres side. The busy connection goes back to the pool, the
next query on it queues
behind the still running statement and times out too, which is the cascade
seen in the field.
### Fixes
- **Demand aware timeout**: new `TimeoutOnPendingDemand` operator, used for
streamed Flux results.
The timer only runs while the downstream has pending demand and restarts
on each delivered row.
A genuine stall while rows are awaited still raises `TimeoutException`.
Mono results are unchanged.
- **Cancel on timeout**: on `TimeoutException`, `PostgresExecutor` sends a
r2dbc-postgresql
`cancelRequest()` before releasing the connection, so it is reusable
within milliseconds.
- **Library upgrade**: jOOQ 3.20.5 -> 3.20.20 and r2dbc-postgresql 1.0.7 ->
1.1.3, which ship the
fixes for cancelled queries poisoning pooled connections (jOOQ#20112,
r2dbc-postgresql#661).
### Tests (TDD, reproducing tests committed before each fix)
- `PostgresReIndexingIntegrationTest`: the reported webadmin scenario on a
Guice Postgres server.
- `PostgresExecutorTimeoutTest`: slow consumer must not time out, a stalled
query must, and a
single-connection pool is usable right after a timeout (60s `pg_sleep`).
- `TimeoutOnPendingDemandTest`: 10 virtual-time cases for the operator.
Validated on the Postgres backend, mailbox, blob, data, JMAP data,
event-bus, task and vault
module suites, and the full Postgres webadmin integration module.
### Note for reviewers
As discussed on the list, paginating the mailbox listing (per-page queries)
would additionally avoid
holding a pooled connection during slow consumption, and would preserve the
previous accidental
"deadlock breaker" behaviour of the timeout under pool exhaustion. This PR
keeps streaming and can
be combined with that follow-up.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]