On Sun, 20 Sep 2026, Richard Henderson wrote:

> On 9/20/26 01:46, Mikulas Patocka wrote:
> > I'm experiencing random deadlocks when running heavily multithreaded
> > workload in qemu-sh4 userspace emulation. This patch fixes them.
> > 
> > The sh4 architecture doesn't support multiprocessing, the processor
> > doesn't have atomic instructions and it uses a technique known as gUSA to
> > provide atomicity guarantees w.r.t. signals or thread scheduling (see the
> > commit 3b894b699c9a for a brief description of gUSA).
> > 
> > The function decode_gusa attempts to recognize several well-known gUSA
> > regions and turn them into atomic instructions. When it fails to
> > recognize a known gUSA pattern, it generates a call to helper_exclusive.
> > Qemu will attempt to stop all the other threads and execute a gUSA region
> > exclusively, so that it can't race with anything.
> > 
> > The problem is in the function cpu_exec_step_atomic - this function stops
> > all other threads with start_exclusive(), then it executes one
> > instruction and then it releases all other threads with end_exclusive().
> > This works for all the architectures except sh4. On sh4, executing one
> > instruction atomically is not enough - we must execute the full gUSA
> > region while the other threads are stopped.
> > 
> > This patch fixes cpu_exec_step_atomic - it adds two new per-architecture
> > functions: is_uninterruptible and revert_uninterruptible.
> > 
> > is_uninterruptible returns true if we are in a gUSA region and we should
> > continue executing code while the other threads are stopped.
> > 
> > If we got TB_EXIT_REQUESTED, we must stop executing code - in this case,
> > we call the function revert_uninterruptible that rolls back PC to the
> > beginning of the gUSA region.
> > 
> > Cc:[email protected]
> > Signed-off-by: Mikulas Patocka<[email protected]>
> > 
> > ---
> >   accel/tcg/cpu-exec.c        |   15 +++++++++++++++
> >   include/accel/tcg/cpu-ops.h |   20 ++++++++++++++++++++
> >   target/sh4/cpu.c            |   25 ++++++++++++++++++++++++-
> >   3 files changed, 59 insertions(+), 1 deletion(-)
> 
> I think this should be integrated into the sh4 translator instead.  We should
> not be single-stepping through the atomic region, but generate one TB that
> implements the entire region.
> 
> This may require some coordination with the translator loop.
> 
> 
> r~

Hi

I looked at the sh4 translator and it seems that it already tries to make 
sure that the full gUSA region is translated into one TB - i.e. there is 
"ctx->base.max_insns = max_insns" in sh4_tr_init_disas_context - that will 
override the value "1" that is supplied by cpu_exec_step_atomic. The 
comments suggest that the author is aware of the fact that the gUSA region 
must be completed atomically.

So, the code that single-steps through the gUSA region is not needed.

It seems that the misbehavior is caused by the fact that if we exit from 
the gUSA TB early, we execute the rest of the region in non-exclusive 
context.

I simplified the patch, so that it adds just one method - 
revert_uninterruptible. It tests whether we are in the unfinished gUSA 
region, and if we are, it reverts PC and SP back to the beginning.

Here I'm sending the updated patch.

Mikulas



From: Mikulas Patocka <[email protected]>

I'm experiencing random deadlocks when running heavily multithreaded
workload in qemu-sh4 userspace emulation. This patch fixes them.

The sh4 architecture doesn't support multiprocessing, the processor
doesn't have atomic instructions and it uses a technique known as gUSA to
provide atomicity guarantees w.r.t. signals or thread scheduling (see the
commit 3b894b699c9a for a brief description of gUSA).

The function decode_gusa attempts to recognize several well-known gUSA
regions and turn them into atomic instructions. When it fails to
recognize a known gUSA pattern, it generates a call to helper_exclusive.
Qemu will attempt to stop all the other threads and execute a gUSA region
exclusively, so that it can't race with anything.

The problem is in the function cpu_exec_step_atomic - this function stops
all other threads with start_exclusive(), then it executes one
TB and then it releases all other threads with end_exclusive(). If the TB
executing the gUSA region exited early, the exclusive lock is dropped and
the execution continues without holding it - this is the root cause for
this bug.

This patch fixes cpu_exec_step_atomic - it adds a new per-architecture
function: revert_uninterruptible. Its implementation
superh_cpu_revert_uninterruptible tests if we are in an unfinished gUSA
region and rolls back the PC to the beginning of the region.

Cc: [email protected]
Signed-off-by: Mikulas Patocka <[email protected]>

---
 accel/tcg/cpu-exec.c        |    5 +++++
 include/accel/tcg/cpu-ops.h |   10 ++++++++++
 target/sh4/cpu.c            |   25 ++++++++++++++++++++++++-
 3 files changed, 39 insertions(+), 1 deletion(-)

Index: qemu/accel/tcg/cpu-exec.c
===================================================================
--- qemu.orig/accel/tcg/cpu-exec.c      2026-09-22 18:31:48.000000000 +0200
+++ qemu/accel/tcg/cpu-exec.c   2026-09-22 18:31:48.000000000 +0200
@@ -590,6 +590,11 @@ void cpu_exec_step_atomic(CPUState *cpu)
         cpu_exec_longjmp_cleanup(cpu);
     }
 
+#ifdef CONFIG_USER_ONLY
+    if (cpu->cc->tcg_ops->revert_uninterruptible)
+        cpu->cc->tcg_ops->revert_uninterruptible(cpu);
+#endif
+
     /*
      * As we start the exclusive region before codegen we must still
      * be in the region if we longjump out of either the codegen or
Index: qemu/include/accel/tcg/cpu-ops.h
===================================================================
--- qemu.orig/include/accel/tcg/cpu-ops.h       2026-09-22 18:31:48.000000000 
+0200
+++ qemu/include/accel/tcg/cpu-ops.h    2026-09-22 18:31:48.000000000 +0200
@@ -168,6 +168,16 @@ struct TCGCPUOps {
      * @addr: tagged guest address
      */
     vaddr (*untagged_addr)(CPUState *cs, vaddr addr);
+
+    /**
+     * revert_uninterruptible:
+     * @cpu: cpu context
+     *
+     * This function is called if cpu_exec_step_atomic needs to exit. It
+     * tests if we are in the gUSA region and rolls back PC to the
+     * beginning of it.
+     */
+    void (*revert_uninterruptible)(CPUState *cs);
 #else
     /** @do_interrupt: Callback for interrupt handling.  */
     void (*do_interrupt)(CPUState *cpu);
Index: qemu/target/sh4/cpu.c
===================================================================
--- qemu.orig/target/sh4/cpu.c  2026-09-22 18:31:48.000000000 +0200
+++ qemu/target/sh4/cpu.c       2026-09-22 18:31:48.000000000 +0200
@@ -92,6 +92,27 @@ static void superh_restore_state_to_opc(
      */
 }
 
+#ifdef CONFIG_USER_ONLY
+static void superh_cpu_revert_uninterruptible(CPUState *cs)
+{
+    SuperHCPU *cpu = SUPERH_CPU(cs);
+    /*
+     * If we are interrupted in the middle of the gUSA region, we must
+     * roll-back PC to the beginning of the region. Continuing halfway
+     * through the region would break atomicity guarantees.
+     *
+     * If we are interrupted after the final write instruction (i.e.
+     * cpu->env.pc == cpu->env.gregs[0]), we must not roll-back, because
+     * the atomic write is already committed in the memory.
+     */
+    if (cpu->env.gregs[15] >= -128u && cpu->env.pc < cpu->env.gregs[0]) {
+        cpu->env.pc = cpu->env.gregs[0] + cpu->env.gregs[15] - 2;
+        cpu->env.gregs[15] = cpu->env.gregs[1];
+        cpu->env.flags &= ~TB_FLAG_ENVFLAGS_MASK;
+    }
+}
+#endif /* CONFIG_USER_ONLY */
+
 #ifndef CONFIG_USER_ONLY
 static bool superh_io_recompile_replay_branch(CPUState *cs,
                                               const TranslationBlock *tb)
@@ -308,7 +329,9 @@ static const TCGCPUOps superh_tcg_ops =
     .restore_state_to_opc = superh_restore_state_to_opc,
     .mmu_index = sh4_cpu_mmu_index,
 
-#ifndef CONFIG_USER_ONLY
+#ifdef CONFIG_USER_ONLY
+    .revert_uninterruptible = superh_cpu_revert_uninterruptible,
+#else
     .tlb_fill = superh_cpu_tlb_fill,
     .pointer_wrap = cpu_pointer_wrap_notreached,
     .cpu_exec_interrupt = superh_cpu_exec_interrupt,


Reply via email to