This series extends Mathieu Desnoyers's hazard pointer implementation [1] with a shared scan path for concurrent hazptr_synchronize() callers, and adapts the lockdep dynamic-key hashlist use case from Boqun Feng's 2025 shazptr series [2] to the current hazptr API.
The series also adds rcuscale support, torture coverage, and litmus tests for the hazptr implementation. The lockdep conversion replaces the expedited RCU wait in lockdep_unregister_key() with hazptr_synchronize(). This limits the wait to hazard pointers protecting the target hash bucket instead of waiting for a system-wide expedited RCU grace period. [1] https://lore.kernel.org/all/[email protected]/ [2] https://lore.kernel.org/lkml/[email protected]/ Performance data ================ The measurements below were collected on an ARM64 KVM guest running on an ARM64 server (96 vCPUs, 8 GB RAM), unless noted otherwise. The kernel is based on Paul McKenney's -rcu tree "dev" branch at commit d21906b0aa1e ("doc: Document additional RCU task-stall dump information") with the full hazptr patch series applied. CONFIG_PREEMPT=y, CONFIG_PREEMPT_RCU=y, CONFIG_NR_CPUS=256. Configuration is saved alongside each result set. Following Paul's suggestion, the reader-side refscale numbers are reported separately for CONFIG_PROVE_LOCKING=n and CONFIG_PROVE_LOCKING=y, since lockdep instrumentation significantly affects the measured RCU and SRCU reader-side costs but has little effect on the hazptr reader path in this setup. Lockdep workload -- tc qdisc mq x100, CONFIG_PROVE_LOCKING=y, 96 background hazptr readers: function-call IPIs (100 tc qdisc operations): 183 For reference, insmod/rmmod x10 in the same setup generated 4930 function-call IPIs. The lockdep_unregister_key() path is the motivating use case for this series: it replaces a system-wide synchronize_rcu_expedited() call with a hazptr_synchronize() scoped to the per-key hash bucket. Writer-side effective per-GP time (rcuscale, nwriters=16, nreaders=0, one run per CPU count; computed as total test duration divided by the total number of grace periods across all writers): hazptr RCU Tree SRCU CPUs per-GP per-GP per-GP ----------------------------------------------------------- 24 498 us 863 us 632 us 96 493 us 865 us 599 us 128 493 us 873 us 789 us 256 500 us 1162 us 476 us At 96 CPUs the measurements are from the comparison blocks in the same guest boot (order: hazptr nw=16, hazptr nw=1, hazptr nw=16, rcu nw=1, rcu nw=16, srcu nw=1, srcu nw=16). The hazptr scalability row uses the first hazptr nw=16 block, which is consistent with the single per-CPU runs at 24, 128, and 256 CPUs. RCU and Tree SRCU use their standard implementations. Hazptr effective per-GP time stays nearly constant from 24 to 256 CPUs, remaining within 493-500 us, consistent with the shared-scan kthread amortising the scan across concurrent synchronize callers. At 96 CPUs, the effective per-GP times are approximately 865 us for RCU, 599 us for SRCU, and 493 us for hazptr with 16 concurrent callers. For comparison, the nwriters=1, nreaders=0 case measured about 19 us per grace period in this run at 96 CPUs. Reader-side overhead (refscale, nreaders=-1 for 75% of 96 online CPUs, nruns=5, median of 5 runs per scale type): The lockdep=n and lockdep=y refscale numbers were collected in separate QEMU boots from kernels built from the same source tree with PROVE_LOCKING toggled. The first hazptr block in each boot is used. All 96 vCPUs were visible to the guest (CONFIG_NR_CPUS=256). PROVE_LOCKING=n PROVE_LOCKING=y hazptr 26.1 ns 25.6 ns RCU 5.0 ns 165.5 ns SRCU 38.8 ns 192.3 ns Hazptr reader-side overhead is nearly unchanged with and without CONFIG_PROVE_LOCKING in this setup, while RCU and SRCU show substantially higher costs when lockdep is enabled. This is why the refscale results are reported separately for the two configurations. Hazptr torture regression (CONFIG_PROVE_LOCKING=y): basic, lockdep+wq_churn, and SLOWPATH PASS at 96 CPUs; CPU sweep (8/16/32/64/128/256) also PASS with no lockdep warnings. Additional x86 server testing with Lian Wang is planned(maybe after LPC). Test reproducibility --------------------- Across five refscale runs, the standard deviation was 0.16 ns for hazptr and 0.12 ns for RCU with PROVE_LOCKING=n. Rcuscale data points are based on a single test run per CPU count, but each run covers thousands of individual grace periods and the hazptr measurements remain within 493-500us across the tested CPU counts. The lockdep workload measurement uses 100 tc qdisc operations in a single QEMU boot; repeated runs agree to within ~10 IPIs on this server. Changes since v2 ================ - v2: https://lore.kernel.org/all/[email protected]/ Only cover letter changes; no code changes since v2. - Fixed attribution: the base hazptr implementation is Mathieu Desnoyers's work. - Split the refscale reader-overhead table into separate PROVE_LOCKING=n and PROVE_LOCKING=y columns, following Paul's suggestion, so that the lockdep impact on RCU and SRCU fast paths is visible rather than folded into a single number. - Replaced the rcuscale "per-writer latency" column with per-grace-period values. The new metric (total test duration divided by total grace periods) is more directly interpretable when comparing 1-vs-16 concurrent synchronize callers. The nw=16 comparison now uses the same rcuscale test block for hazptr, RCU, and SRCU at each CPU count. The nw=1 baseline is provided for reference so that the reader can see the single- writer cost and the per-GP cost under 16 concurrent callers side by side. - Added test-condition and reproducibility notes (kernel commit, PREEMPT model, visible CPU count, boot ordering, standard deviation, sample sizes). - Removed the unreviewed srcua scale-type patch from this series; it will be posted separately. Changes since RFC/WIP ===================== - RFC/WIP: https://lore.kernel.org/all/[email protected]/ - Split the original 4-patch RFC/WIP into smaller commits covering shared scanning, correctness, API support, lockdep, scaling, and torture testing. - Incorporated Boqun Feng's review feedback: use a Bloom filter to avoid per-waiter allocation, add scoped_guard() support, and add a debug option to force the hazptr acquire slow path. - Fixed scan ordering around backup-slot promotion by scanning all per-CPU slots before the overflow lists, with a separate overflow-list phase. - Simplified the scan cycle to flip first and drain only the old wildcard generation, with herd7-verified LKMM tests for both the in-flight and resolved publication cases. - Extended rcuscale and hazptrtorture coverage, added a selftest script for the torture configurations, and fixed the hazptr_release() kernel-doc. Kunwu Chan (15): hazptr: add shared scan kthread hazptr: use Bloom filter for shared scan waiters hazptr: scan all per-CPU slots before overflow lists hazptr: add scoped_guard() support hazptr: add debug option to force the acquire slow path hazptr: elide redundant first drain pass Documentation/litmus-tests: add hazptr wildcard-flip escape test locking/lockdep: use hazptr to wait for dynamic key lookups rcuscale: add hazptr scale type hazptr: fix kernel-doc of hazptr_release() Documentation/litmus-tests: add hazptr acquire-before-scan test hazptrtorture: add slowpath and lockdep scenarios hazptrtorture: add READERS4 and READERS0 torture configs hazptrtorture: add 128- and 256-CPU configs selftests/rcutorture: add hazptr torture test script Documentation/litmus-tests/README | 13 + .../hazptr/hazptr-acquire-before-scan.litmus | 45 +++ .../hazptr/hazptr-wildcard-flip-escape.litmus | 46 +++ include/linux/hazptr.h | 56 ++- kernel/hazptr.c | 336 +++++++++++++++++- kernel/locking/lockdep.c | 25 +- kernel/rcu/Kconfig.debug | 10 + kernel/rcu/hazptrtorture.c | 57 ++- kernel/rcu/rcuscale.c | 70 +++- .../selftests/rcutorture/bin/hazptr.sh | 146 ++++++++ .../rcutorture/configs/hazptr/CFLIST | 6 + .../rcutorture/configs/hazptr/CPU128 | 16 + .../rcutorture/configs/hazptr/CPU128.boot | 1 + .../rcutorture/configs/hazptr/CPU256 | 16 + .../rcutorture/configs/hazptr/CPU256.boot | 1 + .../rcutorture/configs/hazptr/LOCKDEP | 17 + .../rcutorture/configs/hazptr/LOCKDEP.boot | 1 + .../rcutorture/configs/hazptr/READERS0 | 16 + .../rcutorture/configs/hazptr/READERS0.boot | 2 + .../rcutorture/configs/hazptr/READERS4 | 16 + .../rcutorture/configs/hazptr/READERS4.boot | 2 + .../rcutorture/configs/hazptr/SLOWPATH | 16 + 22 files changed, 881 insertions(+), 33 deletions(-) create mode 100644 Documentation/litmus-tests/hazptr/hazptr-acquire-before-scan.litmus create mode 100644 Documentation/litmus-tests/hazptr/hazptr-wildcard-flip-escape.litmus create mode 100755 tools/testing/selftests/rcutorture/bin/hazptr.sh create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU128.boot create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/CPU256.boot create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/LOCKDEP.boot create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS0.boot create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4 create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/READERS4.boot create mode 100644 tools/testing/selftests/rcutorture/configs/hazptr/SLOWPATH base-commit: d21906b0aa1e9573cdb5e7acaca44966b9d1dcd2 -- 2.43.0

