andygrove commented on issue #6124:
URL: 
https://github.com/apache/datafusion-comet/issues/6124#issuecomment-5861059160

   The worker-starvation explanation in the last paragraph holds up in a 
standalone test. Only Tokio workers drive the I/O and timer driver, so an I/O 
future polled from a Spark task thread's `block_on` gets no wake-ups while 
every worker is busy. With tokio 1.53 and one worker stuck in a 2 s poll, a 10 
ms sleep in another thread's `block_on` took 1.9 s, and so did a socket read 
whose peer wrote after 300 ms. With two workers both busy it was the same.
   
   Two things make that more likely than the executor-cores sizing suggests. On 
a standalone cluster without `spark.executor.cores` the runtime has a single 
worker (#6292). And several JVM calls block a worker from inside a poll 
(#6293), including the S3 credential provider, which is called on every 
request. Running the failing query with `COMET_WORKER_THREADS` set to the 
executor's core count would show whether this is what you're hitting.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to