Identical behavior for 4.1.8

On Monday, April 20, 2026 at 1:49:22 PM UTC-4 Noam Bernstein wrote:

> I'm testing 4.1.8 now
>
> On Monday, April 20, 2026 at 1:44:35 PM UTC-4 Noam Bernstein wrote:
>
>> Specifically, this is happening on 32 core nodes, and it looks like the 
>> second syntax refuses to see the second node.  
>>
>> The first syntax shows one task bound to each physical core on one node, 
>> then one to each physical core on the second node. 
>>
>> The second works up to 8 tasks, and each are bound to 4 cores on the 
>> first node, but as soon as I exceed that it complains, perhaps because it's 
>> trying to bind the tasks that are supposed to be on the 2nd node to the 1st 
>> node's cores.
>>
>> This is openmpi 4.1.4, on Rocky 8, BTW
>>
>> On Monday, April 20, 2026 at 12:56:56 PM UTC-4 Noam Bernstein wrote:
>>
>>> Can anyone tell me why this works (i.e. I have 64 cores available)
>>>      mpiexec -n 64  --report-bindings --bind-to core --rank-by core 
>>> --map-by core --mca pml ucx --mca osc ucx --mca btl ^vader,tcp,openib exec
>>> with the expected bindings while this (should also take 16 * 4 = 64 
>>> cores)
>>>       mpiexec -n 16  --report-bindings --bind-to core --rank-by core 
>>> --map-by slot:PE=4 --mca pml ucx --mca osc ucx --mca btl ^vader,tcp,openib 
>>> exec
>>>
>>> Fails with:
>>>
>>> A request was made to bind to that would result in binding more
>>> processes than cpus on a resource:
>>>
>>>    Bind to:     CORE
>>>    Node:        compute-7-4
>>>    #processes:  2
>>>    #cpus:       1
>>>
>>> You can override this protection by adding the "overload-allowed"
>>> option to your binding directive.
>>>
>>

To unsubscribe from this group and stop receiving emails from it, send an email 
to [email protected].

Reply via email to