On 9/23/26 12:58, Mikulas Patocka wrote:
So, after putting debugging prints into the code, I think I have found
out what happens there:

* decode_gusa calls gen_restart_exclusive
* gen_restart_exclusive generates code that sets TB_FLAG_GUSA_EXCLUSIVE
   and generates a call to helper_exclusive

* when the code is executed, TB_FLAG_GUSA_EXCLUSIVE is set
* helper_exclusive calls cpu_loop_exit_atomic, this makes cpu_exec exit
   with EXCP_ATOMIC
* we go to cpu_loop, we execute cpu_exec_step_atomic
* suppose that exit request is set, cpu_exec_step_atomic does nothing, it
   leaves the CPU in the same state as it was before
* we go back to cpu_loop
* suppose that no signal is delivered, so the gUSA is not rewound
* cpu_loop goes to cpu_exec
* there is one difference - now, TB_FLAG_GUSA_EXCLUSIVE is set and it was
   clear before - so cpu_exec will not use the TB that calls
   helper_exclusive, it will instead use the TB that performs the atomic
   operation (both of these TBs have the same PC, they only differ in
   flags)
* the TB that performs the atomic operation is executed inside cpu_exec
   => race condition

So, I fixed this by clearing TB_FLAG_GUSA_EXCLUSIVE after
cpu_exec_step_atomic - to make sure that cpu_exec will find the TB that
calls helper_exclusive and not the TB that performs the atomic operation.

I tested it and the deadlock is gone. Does this seem reasonable? Or, do
you think that we should fix it somewhere else?

Thanks for the excellent analysis.

While I believe your fix is correct, I think the problem might be more general. I think that the condition "running under atomic step" should be promoted to a general cflag. I also think that the exit check could be suppressed for the single-step. That way we go round the loop fewer times.

I'll prepare a patch.


r~

Reply via email to