From: Bradley Morgan <[email protected]>

In cleanup_srcu_struct(), commit 78a38cbf6f20 ("srcu: Queue sdp->work when
the delay timer is successfully deleted") added logic to queue sdp->work if
timer_delete_sync(&sdp->delay_work) successfully canceled a pending timer
while callbacks remained pending on sdp->srcu_cblist. However, it wrapped
this check in a WARN_ON() under the assumption that callers invoking
srcu_barrier() prior to cleanup_srcu_struct() would prevent the warning
from triggering.

This warning can be spuriously triggered during valid teardown paths where
srcu_barrier() is properly invoked before cleanup_srcu_struct(), such as
when releasing blk-mq tag sets:

WARNING: kernel/rcu/srcutree.c:707 at cleanup_srcu_struct+0x3d6/0x8b0
kernel/rcu/srcutree.c:706
Call Trace:
 <TASK>
 blk_mq_free_tag_set+0x617/0x790 block/blk-mq.c:4976
 scsi_mq_free_tags+0x16/0x30 drivers/scsi/scsi_lib.c:2167
 scsi_remove_host+0x243/0x730 drivers/scsi/hosts.c:193
 uas_disconnect+0x135/0x3e0 drivers/usb/storage/uas.c:1236
 usb_unbind_interface+0x295/0x9f0 drivers/usb/core/driver.c:461
 device_release_driver_internal+0x4f5/0x880 drivers/base/dd.c:1372
 bus_remove_device+0x444/0x560 drivers/base/bus.c:664
 device_del+0x524/0x8f0 drivers/base/core.c:3965
 usb_disconnect+0x346/0x9a0 drivers/usb/core/hub.c:2350
 hub_event+0x1bbb/0x4d30 drivers/usb/core/hub.c:5966
 process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3479
 worker_thread+0xa47/0xfb0 kernel/workqueue.c:3560
 kthread+0x38b/0x480 kernel/kthread.c:436
 ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
 </TASK>

This false positive occurs due to a race between callback invocation and
the teardown thread. First, sdp->delay_work can legitimately remain armed:
when an SRCU grace period completes, srcu_gp_end() arms sdp->delay_work
with a delay, and if callbacks are later scheduled with zero delay via
srcu_schedule_cbs_sdp(sdp, 0), queue_work_on() queues sdp->work directly
without deleting sdp->delay_work. Second, in srcu_invoke_callbacks(),
callbacks (including the one queued by srcu_barrier()) are extracted into a
local list and executed, but sdp->srcu_cblist length is only decremented
via rcu_segcblist_add_len(&sdp->srcu_cblist, -len) after all callbacks
finish executing. When srcu_barrier_cb() executes, it wakes the waiting
srcu_barrier() thread, which proceeds immediately to cleanup_srcu_struct().
At that point, the worker thread is still executing callbacks and has not
yet decremented the callback count, so timer_delete_sync(&sdp->delay_work)
returns 1 and rcu_segcblist_n_cbs(&sdp->srcu_cblist) is non-zero, firing
the WARN_ON(). Immediately thereafter, flush_work(&sdp->work) waits for the
worker to finish, and the subsequent check on rcu_segcblist_n_cbs()
correctly sees no remaining callbacks.

Because WARN_ON() must not be used for conditions that can legitimately
happen, and pr_err() should be used instead if an actual error needs to be
reported (which is not applicable here as this is normal recovery
behavior), remove the WARN_ON() check and its accompanying comment. Keep
the queue_work_on() recovery logic so that flush_work() properly waits for
remaining callbacks to complete. Genuine callback leaks remain caught by
the subsequent authoritative
WARN_ON(rcu_segcblist_n_cbs(&sdp->srcu_cblist)) check performed after work
has been flushed.

Fixes: 78a38cbf6f20 ("srcu: Queue sdp->work when the delay timer is 
successfully deleted")
Assisted-by: Gemini:gemini-3.8-flash Gemini:gemini-3.1-pro-preview syzbot
Reported-by: [email protected]
Closes: https://syzkaller.appspot.com/bug?extid=02b37e31e64ea5cb6d29
Link: 
https://syzkaller.appspot.com/ai_job?id=5a402d0b-b407-49a6-b0f3-b06197c1398f
Signed-off-by: Bradley Morgan <[email protected]>

---
diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c
index ed204b3f4..7f30a5587 100644
--- a/kernel/rcu/srcutree.c
+++ b/kernel/rcu/srcutree.c
@@ -701,10 +701,8 @@ void cleanup_srcu_struct(struct srcu_struct *ssp)
        for_each_possible_cpu(cpu) {
                struct srcu_data *sdp = per_cpu_ptr(ssp->sda, cpu);
 
-               // Call srcu_barrier() before this cleanup_srcu_struct()
-               // to avoid triggering this WARN_ON().
-               if (WARN_ON(timer_delete_sync(&sdp->delay_work) &&
-                           rcu_segcblist_n_cbs(&sdp->srcu_cblist)) &&
+               if (timer_delete_sync(&sdp->delay_work) &&
+                   rcu_segcblist_n_cbs(&sdp->srcu_cblist) &&
                    rcu_cpu_beenfullyonline(sdp->cpu))
                        queue_work_on(sdp->cpu, rcu_gp_wq, &sdp->work);
                flush_work(&sdp->work);


base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
-- 
See https://goo.gle/syzbot-ai-patches for information about AI-generated 
patches.
The person who has signed off on the patch is responsible for
addressing comments.
syzbot engineers can be reached at [email protected].

Reply via email to