On 9/20/26 01:46, Mikulas Patocka wrote:
I'm experiencing random deadlocks when running heavily multithreaded
workload in qemu-sh4 userspace emulation. This patch fixes them.

The sh4 architecture doesn't support multiprocessing, the processor
doesn't have atomic instructions and it uses a technique known as gUSA to
provide atomicity guarantees w.r.t. signals or thread scheduling (see the
commit 3b894b699c9a for a brief description of gUSA).

The function decode_gusa attempts to recognize several well-known gUSA
regions and turn them into atomic instructions. When it fails to
recognize a known gUSA pattern, it generates a call to helper_exclusive.
Qemu will attempt to stop all the other threads and execute a gUSA region
exclusively, so that it can't race with anything.

The problem is in the function cpu_exec_step_atomic - this function stops
all other threads with start_exclusive(), then it executes one
instruction and then it releases all other threads with end_exclusive().
This works for all the architectures except sh4. On sh4, executing one
instruction atomically is not enough - we must execute the full gUSA
region while the other threads are stopped.

This patch fixes cpu_exec_step_atomic - it adds two new per-architecture
functions: is_uninterruptible and revert_uninterruptible.

is_uninterruptible returns true if we are in a gUSA region and we should
continue executing code while the other threads are stopped.

If we got TB_EXIT_REQUESTED, we must stop executing code - in this case,
we call the function revert_uninterruptible that rolls back PC to the
beginning of the gUSA region.

Cc:[email protected]
Signed-off-by: Mikulas Patocka<[email protected]>

---
  accel/tcg/cpu-exec.c        |   15 +++++++++++++++
  include/accel/tcg/cpu-ops.h |   20 ++++++++++++++++++++
  target/sh4/cpu.c            |   25 ++++++++++++++++++++++++-
  3 files changed, 59 insertions(+), 1 deletion(-)

I think this should be integrated into the sh4 translator instead. We should not be single-stepping through the atomic region, but generate one TB that implements the entire region.

This may require some coordination with the translator loop.


r~

Reply via email to