apport information

** Tags added: apport-collected cloud-image

** Description changed:

  # linux-azure: kernel panic in rcu_do_batch (NULL callback) — regression
  between 6.8 and 6.14
  
  **Package:** linux-azure (Ubuntu 24.04.3 LTS, noble)
  **Regression:** clean on 6.8.0-1018-azure; panics from 6.14.0-1017-azure 
onward
  **Still present on:** 6.17.0-1008 / -1010 / -1011 / -1013 / -1015 / -1018 / 
-1021 / -1022
  **Unverified on:** 7.0.0-1014-azure (too little uptime to conclude — see 
Current state)
  **Scope:** 62 panics across 10 VMs in two environments, plus a 3-VM control 
on the same
  kernel that has never panicked
  **Evidence current as of:** 2026-09-29
  
  ## Summary
  
  Two Azure Docker Swarm clusters (5 VMs each) began panicking after being 
rebooted into
  `6.14.0-1017-azure`. The same machines had previously run **379–397 
consecutive days without
  a single panic** on `6.8.0-1018-azure`.
  
  A vmcore captured 2026-09-26 shows RCU invoking a callback whose `rcu_head` 
points at the
  base of a freed, VMAP'd kernel stack, with maple-tree VMA teardown frames 
(`mas_split`,
  `mas_topiary_replace` → `call_rcu`) in the remnants of that stack.
  
  A third cluster running the **same kernel on identical hardware has zero 
panics in 275
  days**. The distinguishing variable we can find is query load against an 
mmap-heavy
  workload — detailed under Trigger below.
  
  ## Regression window — the key evidence
  
  Serial console history goes back to VM creation in September 2023. Zero 
panics in ~27
  months, then panics immediately on 6.14:
  
  | VM | Kernel | Booted | Uptime | Outcome |
  |---|---|---|---|---|
  | vm-qat-3 | 5.15.0-1045-azure | 2023-09-11 | 351.9 d | clean |
  | vm-qat-3 | 6.8.0-1018-azure | 2024-12-12 | **379.2 d** | clean |
  | vm-qat-3 | **6.14.0-1017-azure** | 2025-12-26 | 0.3 d | **PANIC** |
  | vm-qat-4 | 6.8.0-1018-azure | 2024-12-12 | **379.2 d** | clean |
  | vm-qat-4 | **6.14.0-1017-azure** | 2025-12-26 | 16.3 d | **PANIC** |
  | vm-prd-1 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
  | vm-prd-2 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
  | vm-prd-3 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
  | vm-prd-4 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
  | vm-prd-0 | 6.8.0-1018-azure | 2024-12-12 | 343.7 d | clean |
  
  **Seven VMs ran roughly a year each on 6.8.0-1018 with no panic.** 
Production's first panic
  is 2026-01-13, immediately after its own move to 6.14.0-1017.
  
  These hosts ran 6.8.0-1018 continuously and jumped **straight from 6.8 to 
6.14** in a single
  reboot; 6.11 and 6.14.0-1012/-1013/-1014 were installed by 
unattended-upgrades but never
  booted on any affected host. **The bisect window is therefore 6.8 → 6.14**, 
not narrowly
  6.14.0-1017.
  
  **The distro release is not implicated.** These hosts were upgraded 20.04 → 
22.04 → 24.04 on
  2024-09-20, then ran **462 days on 24.04 without a panic** (82 d on 
6.8.0-1015, then 379 d on
  6.8.0-1018) before the 6.14 reboot. The regression tracks the kernel series, 
not the release.
  
  Panic counts since the 6.14 reboot (excluding two deliberate sysrq
  tests):
  
  ```
  QAT   vm-qat-0 14   vm-qat-1 13   vm-qat-2  3   vm-qat-3 10   vm-qat-4  5     
= 45
  PROD  vm-prd-0  4   vm-prd-1  7   vm-prd-2  2   vm-prd-3  2   vm-prd-4  2     
= 17
  ```
  
  ## The control: same kernel, same hardware, zero panics
  
  A third cluster ("latest") booted `6.14.0-1017-azure` on **2025-12-26 — the 
same day as the
  other two** — and has run it continuously ever since:
  
  | VM | Kernel | Booted | Uptime | Panics |
  |---|---|---|---|---|
  | vm-latest-0 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |
  | vm-latest-1 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |
  | vm-latest-2 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |
  
  vm-latest-0 is the **same VM SKU** (Standard_E2as_v5) on the **same CPU** 
(AMD EPYC 7763) as
  every affected host, same distro, same kernel, same service count (33), 
comparable container
  churn. It runs the same Prometheus workload off the same kind of CIFS share.
  
  **The panics are therefore not deterministic per kernel version.** Something 
workload-dependent
  gates them. The same is visible within the affected clusters: vm-prd-2 ran 
6.14.0-1017 for
  123.3 days clean, and vm-qat-2/-3 are currently at 3 weeks and 2.5 weeks on 
6.17.0-1022
  without incident.
  
  ## Trigger: query load against an mmap-heavy workload
  
  The workload is Prometheus with its TSDB on a CIFS mount (Azure Files, 
`cache=strict`).
  Prometheus mmaps block chunk files to serve range queries, so query volume 
drives VMA
  create/destroy churn — which allocates and RCU-frees maple-tree nodes, the 
code in the crash
  stack.
  
  Measured query rates, with ingest volume as a control:
  
  | env | prometheus uptime | queries | **queries/hour** | head series | panics 
|
  |---|---|---|---|---|---|
  | prod | 5.2 d | 1,953,493 | **15,692** | 30,733 | 17 |
  | qat | 3.3 d | 30,585 | **384** | 30,666 | 45 |
  | latest | 192.9 d | 115 | **0.02** | 28,247 | **0** |
  
  **Ingest is effectively identical** across all three (28k–31k active series) 
while query rate
  spans six orders of magnitude. The control cluster's Grafana database was 
wiped months ago,
  so nothing queries its Prometheus — 115 queries in 192 days.
  
  Confirmed these are external API queries, not internal rule evaluation:
  `prometheus_rule_evaluations_total` is absent (no recording or alerting 
rules), and
  `prometheus_engine_query_duration_seconds_count` matches
  `prometheus_http_requests_total{handler=~"/api/v1/query.*"}` to within 0.1%.
  
  One tight sequence on a single host, after pinning the workload there on
  2026-09-24 15:14:03:
  
  ```
  2026-09-24 15:14:01   workload starts on host (CIFS mount)
  2026-09-24 16:01:03   Bad rss-counter / non-zero pgtables_bytes      (+47 min)
  2026-09-26 06:01:03   kernel panic                                   (+38 h 
47 m)
  ```
  
  **Important qualification:** query load looks *necessary* (the zero-query 
cluster has zero
  panics) but is **not simply proportional** to panic frequency — prod queries 
41× harder than
  QAT yet panics less often. QAT has substantially heavier container churn from 
CI deploys, and
  the mm precursor below always fires on container-runtime teardown, so a 
second factor is
  likely involved.
  
  ## Panic signatures
  
  Three distinct signatures, all consistent with one memory corruptor:
  
  1. `Kernel panic - not syncing: Fatal exception in interrupt` (most common)
  2. `Kernel panic - not syncing: corrupted stack end detected inside scheduler`
     (`CONFIG_SCHED_STACK_END_CHECK=y`)
  3. `Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000000b` 
(once)
  
  Representative panic (2026-09-26, 6.17.0-1022-azure):
  
  ```
  BUG: kernel NULL pointer dereference, address: 0000000000000000
  #PF: supervisor instruction fetch in kernel mode
  Oops: Oops: 0010 [#1] SMP NOPTI
  CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Kdump: loaded Not tainted 
6.17.0-1022-azure #22-Ubuntu
  RIP: 0010:0x0
  RDI: ffffd1d149880000
  Call Trace:
   <IRQ>
   rcu_do_batch+0x1b4/0x560
   rcu_core+0x117/0x1e0
   rcu_core_si+0xe/0x20
   handle_softirqs+0xe2/0x2f0
   __irq_exit_rcu+0xdd/0x100
   irq_exit_rcu+0xe/0x20
   sysvec_apic_timer_interrupt+0x8a/0xb0
   </IRQ>
   asm_sysvec_apic_timer_interrupt+0x1b/0x20
  RIP: 0010:pv_native_safe_halt+0xb/0x10
  ```
  
  The CPU was idle (`swapper/1`, `pv_native_safe_halt`) with **32.7 hours of 
complete kernel
  log silence** before the panic — the corrupted callback simply sat in the 
list until RCU
  reached it.
  
  ## vmcore analysis
  
  Analysed with `crash 8.0.4` + `linux-image-6.17.0-1022-azure-dbgsym`.
  
  **The bad `rcu_head` is the base of a VMAP'd kernel stack.** The mapped 
region is exactly
  16 KB (`ffffd1d149880000`–`ffffd1d149883fff`) with **unmapped guard pages on 
both sides** —
  `THREAD_SIZE`, and `CONFIG_VMAP_STACK=y`:
  
  ```
  crash> rd ffffd1d14987f000 2
  rd: invalid kernel virtual address: ffffd1d14987f000     <- guard page
  crash> rd ffffd1d149880000 2 ... ffffd1d149883000 2      <- 4 pages mapped
  crash> rd ffffd1d149884000 2
  rd: invalid kernel virtual address: ffffd1d149884000     <- guard page
  crash> kmem ffffd1d149880000
  kmem: invalid kernel virtual address                     <- not a tracked 
allocation
  ```
  
  `vtop` shows it PRESENT|RW|DIRTY|NX. The stack had been freed and zeroed, so 
`rhp->func`
  read as 0.
  
  **CPU 1's callback list at panic:**
  
  ```
  crash> struct rcu_data.cblist rcu_data:1
  [1]: ffff8937bfd323c0
    cblist = {
      head = 0xffff89352a20aa80,
      tails = {0xffff8937bfd32440, 0xffff89352a20aa80, ...},
      len = { counter = 56 }, seglen = {0, 1, 0, 0}, flags = 1
    }
  ```
  
  `R12` in the panic registers equals `rcu_data:1` exactly; `RBX=0x37` (55) is 
the loop
  position — it died on roughly the 55th of 56 callbacks. The local 
`rcu_cblist` on the stack
  was fully drained (`head=NULL, tail=&self`), so the bad pointer came from the 
previous
  entry's `->next`.
  
  **Remnants on that kernel stack name the subsystem:**
  
  ```
  mas_split+1409
  mas_topiary_replace+3236
  __call_rcu_common.constprop.0+183
  call_rcu+14
  ```
  
  interleaved with userspace VMA bound pairs (e.g. 
`0000796f41a00000`/`0000796f41ceafff`).
  The stack belonged to a task doing mmap/munmap VMA work, freeing maple-tree 
nodes via
  `call_rcu()`.
  
  ## Precursor: mm accounting corruption
  
  Every affected boot shows this in dmesg, always on container-runtime 
address-space teardown,
  hours to days before the panic:
  
  ```
  BUG: Bad rss-counter state mm:ffff8934aa46a0c0 type:MM_ANONPAGES val:3 
Comm:containerd-shim Pid:3802600
  BUG: non-zero pgtables_bytes on freeing mm: 8192
  ```
  
  Also seen with `Comm:runc` (once with `non-zero pgtables_bytes: 49152`).
  
  ## Current state (2026-09-29)
  
  | VM | Kernel | Uptime | Notes |
  |---|---|---|---|
  | vm-qat-0 | 7.0.0-1014 | 3 d 8 h | workload resident |
  | vm-qat-1 | 7.0.0-1014 | 6 d 4 h | |
  | vm-qat-2 | 6.17.0-1022 | 3 w 2 d | |
  | vm-qat-3 | 6.17.0-1022 | 2 w 5 d | |
  | vm-qat-4 | 7.0.0-1014 | 23 h | |
  | prd-0..4 | 6.17.0-1018 / -1022 / 7.0.0-1014 | 4–77 d | |
  | latest-0..2 | 6.14.0-1017 | ~275 d | control, zero panics |
  
  No conclusion should be drawn yet about 7.0.0-1014. Historical intervals 
between panics on
  affected kernels ranged **0.4 – 60.5 days** (median ~10), so current uptimes 
sit well inside
  the range that affected kernels also survived.
  
  ## What we ruled out
  
  - **Not slab corruption.** `slub_debug=FZPU` ran for 17 days with redzone + 
poison + user
    tracking active on all 545 caches (`red_zone=1 on 545 of 545`), across a 
panic, and
    produced **zero** reports. Consistent with the corrupted object being a 
vmalloc'd kernel
    stack rather than slab memory.
  - **Not resource exhaustion.** At panic: ~1 GB of 15 GB used, CPU ~5%, OS 
disk IOPS 0%.
  - **Not the hypervisor.** Azure Resource Health shows no platform event 
coinciding with any
    of the 62 panics. (Separately, vm-qat-4 hit an unrelated Azure 
host-degradation incident on
    2026-09-25 — a storage I/O hang with no panic. Different failure mode, 
excluded from the
    counts.)
  - **Not hardware variation.** Every affected host and the control node 
vm-latest-0 are the
    same SKU (Standard_E2as_v5) on the same CPU (AMD EPYC 7763).
  - **Userspace packages are implausible but not formally excluded** — see 
Limits.
  
  ## Limits of this evidence
  
  Stated explicitly so nothing here is over-read:
  
  - **The 26 Dec reboot changed two things, not one.** A manual `apt upgrade` 
ran on each host
    minutes before its reboot, carrying dmidecode, libpolkit-agent-1-0, 
dmeventd,
    initramfs-tools-core, netplan-generator, libsmartcols1, ssh-import-id and 
six python3
    packages, plus a 24.04.x point release of base-files. So kernel and 
userspace changed
    together. We consider userspace implausible as a cause — none of those 
packages ships
    kernel code, and these hosts boot initrdless via `GRUB_FORCE_PARTUUID`, 
making
    initramfs-tools inert — but the before/after boundary alone does not 
isolate the kernel.
  - **Uptimes are lower bounds.** "Ran N days" is derived from the last 
timestamped line in
    each boot's console output, and "clean" means no panic was printed — a 
silent hang would
    be indistinguishable.
  - **The maple-tree frames are stale stack remnants**, recovered from freed 
memory, not a live
    call trace. They establish what that stack had been doing, not what 
corrupted it.
  - **The trigger correlation is across three environments, not a controlled 
test.** Query rate
    is a proxy for mmap/munmap rate, which we did not measure directly. And it 
is not
    proportional to panic frequency (see Trigger).
  - **6.11.x is untested** on any affected host — it was installed but never 
booted, so the
    bisect window cannot currently be narrowed below 6.8 → 6.14.
  
  A planned change will test the first point directly: moving these hosts to the
  `linux-azure-6.8` GA track runs **current** userspace against the older 
kernel series. If the
  panics stop, the kernel is isolated cleanly. QAT moves first, soaks two 
weeks, then prod.
  
  ## System details
  
  - Ubuntu 24.04.3 LTS (noble)
  - Azure `Standard_E2as_v5` — 2 vCPU, 16 GB, AMD EPYC 7763
  - `CONFIG_VMAP_STACK=y`, `CONFIG_SCHED_STACK_END_CHECK=y`,
    `CONFIG_DEBUG_OBJECTS` **not set**, `CONFIG_KASAN` **not set**
  - Kernel cmdline: `console=tty1 console=ttyS0 earlyprintk=ttyS0 panic=-1` + 
`crashkernel=`
  
  Note: neither `CONFIG_DEBUG_OBJECTS` (which would catch a bad/on-stack 
`rcu_head` directly)
  nor `CONFIG_KASAN` is available in the shipped Azure kernel, so we cannot 
instrument this
  further without a custom build. **A test kernel with `DEBUG_OBJECTS_RCU_HEAD` 
would very
  likely identify the offending `call_rcu()` caller immediately, and we can run 
one** on a host
  that reproduces.
  
  ## Attachments
  
  - `dmesg-20260926.txt` — full kernel ring buffer at panic, extracted by 
makedumpfile
  - `dump.202609260602` — 1.7 GB vmcore (kdump, `-c -d 31`), available on 
request
  
  kdump has since been reconfigured to `-c -d 14`, so the next capture will 
retain free and
  zero pages — the memory a freed-kernel-stack bug most needs, and which `-d 
31` stripped from
  the dump above. We can supply a richer core after the next panic.
+ --- 
+ ProblemType: Bug
+ ApportVersion: 2.28.1-0ubuntu3.8
+ Architecture: amd64
+ AzureImageoffer: 0001-com-ubuntu-server-focal
+ AzureImagepublisher: canonical
+ AzureImagesku: 20_04-lts-gen2
+ AzureImageversion: 20.04.202308310
+ AzureVmsize: Standard_E2as_v5
+ CasperMD5CheckResult: unknown
+ CloudArchitecture: x86_64
+ CloudBuildName: server
+ CloudID: azure
+ CloudName: azure
+ CloudPlatform: azure
+ CloudRegion: eastus
+ CloudSerial: 20230831
+ CloudSubPlatform: seed-dir (/var/lib/waagent)
+ DistroRelease: Ubuntu 24.04
+ Package: linux-azure 7.0.0-1014.14~24.04.1
+ PackageArchitecture: amd64
+ ProcEnviron:
+  LANG=C.UTF-8
+  PATH=(custom, no user)
+  SHELL=/bin/bash
+  TERM=xterm-256color
+  XDG_RUNTIME_DIR=<set>
+ ProcVersionSignature: User Name 7.0.0-1014.14~24.04.1-azure 7.0.14
+ Tags: cloud-image noble
+ Uname: Linux 7.0.0-1014-azure x86_64
+ UpgradeStatus: Upgraded to noble on 2024-09-20 (739 days ago)
+ UserGroups: adm audio cdrom dialout dip docker floppy lxd netdev plugdev sudo 
video
+ _MarkForUpload: True

** Attachment added: "Dependencies.txt"
   
https://bugs.launchpad.net/bugs/2168861/+attachment/6003604/+files/Dependencies.txt

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168861

Title:
  linux-azure: kernel panic in rcu_do_batch (NULL rcu_head on freed VMAP
  stack) — regression between 6.8 and 6.14

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-azure/+bug/2168861/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to