On 9/23/26 12:58, Mikulas Patocka wrote:
So, after putting debugging prints into the code, I think I have found
out what happens there:
* decode_gusa calls gen_restart_exclusive
* gen_restart_exclusive generates code that sets TB_FLAG_GUSA_EXCLUSIVE
and generates a call to helper_exclusive
* when the code is executed, TB_FLAG_GUSA_EXCLUSIVE is set
* helper_exclusive calls cpu_loop_exit_atomic, this makes cpu_exec exit
with EXCP_ATOMIC
* we go to cpu_loop, we execute cpu_exec_step_atomic
* suppose that exit request is set, cpu_exec_step_atomic does nothing, it
leaves the CPU in the same state as it was before
* we go back to cpu_loop
* suppose that no signal is delivered, so the gUSA is not rewound
* cpu_loop goes to cpu_exec
* there is one difference - now, TB_FLAG_GUSA_EXCLUSIVE is set and it was
clear before - so cpu_exec will not use the TB that calls
helper_exclusive, it will instead use the TB that performs the atomic
operation (both of these TBs have the same PC, they only differ in
flags)
* the TB that performs the atomic operation is executed inside cpu_exec
=> race condition
So, I fixed this by clearing TB_FLAG_GUSA_EXCLUSIVE after
cpu_exec_step_atomic - to make sure that cpu_exec will find the TB that
calls helper_exclusive and not the TB that performs the atomic operation.
I tested it and the deadlock is gone. Does this seem reasonable? Or, do
you think that we should fix it somewhere else?
Thanks for the excellent analysis.
While I believe your fix is correct, I think the problem might be more
general. I think that the condition "running under atomic step" should
be promoted to a general cflag. I also think that the exit check could
be suppressed for the single-step. That way we go round the loop fewer
times.
I'll prepare a patch.
r~