Public bug reported:

# linux-azure: kernel panic in rcu_do_batch (NULL callback) — regression
between 6.8 and 6.14

**Package:** linux-azure (Ubuntu 24.04.3 LTS, noble)
**Regression:** clean on 6.8.0-1018-azure; panics from 6.14.0-1017-azure onward
**Still present on:** 6.17.0-1008 / -1010 / -1011 / -1013 / -1015 / -1018 / 
-1021 / -1022
**Unverified on:** 7.0.0-1014-azure (too little uptime to conclude — see 
Current state)
**Scope:** 62 panics across 10 VMs in two environments, plus a 3-VM control on 
the same
kernel that has never panicked
**Evidence current as of:** 2026-09-29

## Summary

Two Azure Docker Swarm clusters (5 VMs each) began panicking after being 
rebooted into
`6.14.0-1017-azure`. The same machines had previously run **379–397 consecutive 
days without
a single panic** on `6.8.0-1018-azure`.

A vmcore captured 2026-09-26 shows RCU invoking a callback whose `rcu_head` 
points at the
base of a freed, VMAP'd kernel stack, with maple-tree VMA teardown frames 
(`mas_split`,
`mas_topiary_replace` → `call_rcu`) in the remnants of that stack.

A third cluster running the **same kernel on identical hardware has zero panics 
in 275
days**. The distinguishing variable we can find is query load against an 
mmap-heavy
workload — detailed under Trigger below.

## Regression window — the key evidence

Serial console history goes back to VM creation in September 2023. Zero panics 
in ~27
months, then panics immediately on 6.14:

| VM | Kernel | Booted | Uptime | Outcome |
|---|---|---|---|---|
| vm-qat-3 | 5.15.0-1045-azure | 2023-09-11 | 351.9 d | clean |
| vm-qat-3 | 6.8.0-1018-azure | 2024-12-12 | **379.2 d** | clean |
| vm-qat-3 | **6.14.0-1017-azure** | 2025-12-26 | 0.3 d | **PANIC** |
| vm-qat-4 | 6.8.0-1018-azure | 2024-12-12 | **379.2 d** | clean |
| vm-qat-4 | **6.14.0-1017-azure** | 2025-12-26 | 16.3 d | **PANIC** |
| vm-prd-1 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
| vm-prd-2 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
| vm-prd-3 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
| vm-prd-4 | 6.8.0-1018-azure | 2024-12-12 | **396.7 d** | clean |
| vm-prd-0 | 6.8.0-1018-azure | 2024-12-12 | 343.7 d | clean |

**Seven VMs ran roughly a year each on 6.8.0-1018 with no panic.** Production's 
first panic
is 2026-01-13, immediately after its own move to 6.14.0-1017.

These hosts ran 6.8.0-1018 continuously and jumped **straight from 6.8 to 
6.14** in a single
reboot; 6.11 and 6.14.0-1012/-1013/-1014 were installed by unattended-upgrades 
but never
booted on any affected host. **The bisect window is therefore 6.8 → 6.14**, not 
narrowly
6.14.0-1017.

**The distro release is not implicated.** These hosts were upgraded 20.04 → 
22.04 → 24.04 on
2024-09-20, then ran **462 days on 24.04 without a panic** (82 d on 6.8.0-1015, 
then 379 d on
6.8.0-1018) before the 6.14 reboot. The regression tracks the kernel series, 
not the release.

Panic counts since the 6.14 reboot (excluding two deliberate sysrq
tests):

```
QAT   vm-qat-0 14   vm-qat-1 13   vm-qat-2  3   vm-qat-3 10   vm-qat-4  5     = 
45
PROD  vm-prd-0  4   vm-prd-1  7   vm-prd-2  2   vm-prd-3  2   vm-prd-4  2     = 
17
```

## The control: same kernel, same hardware, zero panics

A third cluster ("latest") booted `6.14.0-1017-azure` on **2025-12-26 — the 
same day as the
other two** — and has run it continuously ever since:

| VM | Kernel | Booted | Uptime | Panics |
|---|---|---|---|---|
| vm-latest-0 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |
| vm-latest-1 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |
| vm-latest-2 | 6.14.0-1017-azure | 2025-12-26 | ~275 d | **0** |

vm-latest-0 is the **same VM SKU** (Standard_E2as_v5) on the **same CPU** (AMD 
EPYC 7763) as
every affected host, same distro, same kernel, same service count (33), 
comparable container
churn. It runs the same Prometheus workload off the same kind of CIFS share.

**The panics are therefore not deterministic per kernel version.** Something 
workload-dependent
gates them. The same is visible within the affected clusters: vm-prd-2 ran 
6.14.0-1017 for
123.3 days clean, and vm-qat-2/-3 are currently at 3 weeks and 2.5 weeks on 
6.17.0-1022
without incident.

## Trigger: query load against an mmap-heavy workload

The workload is Prometheus with its TSDB on a CIFS mount (Azure Files, 
`cache=strict`).
Prometheus mmaps block chunk files to serve range queries, so query volume 
drives VMA
create/destroy churn — which allocates and RCU-frees maple-tree nodes, the code 
in the crash
stack.

Measured query rates, with ingest volume as a control:

| env | prometheus uptime | queries | **queries/hour** | head series | panics |
|---|---|---|---|---|---|
| prod | 5.2 d | 1,953,493 | **15,692** | 30,733 | 17 |
| qat | 3.3 d | 30,585 | **384** | 30,666 | 45 |
| latest | 192.9 d | 115 | **0.02** | 28,247 | **0** |

**Ingest is effectively identical** across all three (28k–31k active series) 
while query rate
spans six orders of magnitude. The control cluster's Grafana database was wiped 
months ago,
so nothing queries its Prometheus — 115 queries in 192 days.

Confirmed these are external API queries, not internal rule evaluation:
`prometheus_rule_evaluations_total` is absent (no recording or alerting rules), 
and
`prometheus_engine_query_duration_seconds_count` matches
`prometheus_http_requests_total{handler=~"/api/v1/query.*"}` to within 0.1%.

One tight sequence on a single host, after pinning the workload there on
2026-09-24 15:14:03:

```
2026-09-24 15:14:01   workload starts on host (CIFS mount)
2026-09-24 16:01:03   Bad rss-counter / non-zero pgtables_bytes      (+47 min)
2026-09-26 06:01:03   kernel panic                                   (+38 h 47 
m)
```

**Important qualification:** query load looks *necessary* (the zero-query 
cluster has zero
panics) but is **not simply proportional** to panic frequency — prod queries 
41× harder than
QAT yet panics less often. QAT has substantially heavier container churn from 
CI deploys, and
the mm precursor below always fires on container-runtime teardown, so a second 
factor is
likely involved.

## Panic signatures

Three distinct signatures, all consistent with one memory corruptor:

1. `Kernel panic - not syncing: Fatal exception in interrupt` (most common)
2. `Kernel panic - not syncing: corrupted stack end detected inside scheduler`
   (`CONFIG_SCHED_STACK_END_CHECK=y`)
3. `Kernel panic - not syncing: Attempted to kill init! exitcode=0x0000000b` 
(once)

Representative panic (2026-09-26, 6.17.0-1022-azure):

```
BUG: kernel NULL pointer dereference, address: 0000000000000000
#PF: supervisor instruction fetch in kernel mode
Oops: Oops: 0010 [#1] SMP NOPTI
CPU: 1 UID: 0 PID: 0 Comm: swapper/1 Kdump: loaded Not tainted 
6.17.0-1022-azure #22-Ubuntu
RIP: 0010:0x0
RDI: ffffd1d149880000
Call Trace:
 <IRQ>
 rcu_do_batch+0x1b4/0x560
 rcu_core+0x117/0x1e0
 rcu_core_si+0xe/0x20
 handle_softirqs+0xe2/0x2f0
 __irq_exit_rcu+0xdd/0x100
 irq_exit_rcu+0xe/0x20
 sysvec_apic_timer_interrupt+0x8a/0xb0
 </IRQ>
 asm_sysvec_apic_timer_interrupt+0x1b/0x20
RIP: 0010:pv_native_safe_halt+0xb/0x10
```

The CPU was idle (`swapper/1`, `pv_native_safe_halt`) with **32.7 hours of 
complete kernel
log silence** before the panic — the corrupted callback simply sat in the list 
until RCU
reached it.

## vmcore analysis

Analysed with `crash 8.0.4` + `linux-image-6.17.0-1022-azure-dbgsym`.

**The bad `rcu_head` is the base of a VMAP'd kernel stack.** The mapped region 
is exactly
16 KB (`ffffd1d149880000`–`ffffd1d149883fff`) with **unmapped guard pages on 
both sides** —
`THREAD_SIZE`, and `CONFIG_VMAP_STACK=y`:

```
crash> rd ffffd1d14987f000 2
rd: invalid kernel virtual address: ffffd1d14987f000     <- guard page
crash> rd ffffd1d149880000 2 ... ffffd1d149883000 2      <- 4 pages mapped
crash> rd ffffd1d149884000 2
rd: invalid kernel virtual address: ffffd1d149884000     <- guard page
crash> kmem ffffd1d149880000
kmem: invalid kernel virtual address                     <- not a tracked 
allocation
```

`vtop` shows it PRESENT|RW|DIRTY|NX. The stack had been freed and zeroed, so 
`rhp->func`
read as 0.

**CPU 1's callback list at panic:**

```
crash> struct rcu_data.cblist rcu_data:1
[1]: ffff8937bfd323c0
  cblist = {
    head = 0xffff89352a20aa80,
    tails = {0xffff8937bfd32440, 0xffff89352a20aa80, ...},
    len = { counter = 56 }, seglen = {0, 1, 0, 0}, flags = 1
  }
```

`R12` in the panic registers equals `rcu_data:1` exactly; `RBX=0x37` (55) is 
the loop
position — it died on roughly the 55th of 56 callbacks. The local `rcu_cblist` 
on the stack
was fully drained (`head=NULL, tail=&self`), so the bad pointer came from the 
previous
entry's `->next`.

**Remnants on that kernel stack name the subsystem:**

```
mas_split+1409
mas_topiary_replace+3236
__call_rcu_common.constprop.0+183
call_rcu+14
```

interleaved with userspace VMA bound pairs (e.g. 
`0000796f41a00000`/`0000796f41ceafff`).
The stack belonged to a task doing mmap/munmap VMA work, freeing maple-tree 
nodes via
`call_rcu()`.

## Precursor: mm accounting corruption

Every affected boot shows this in dmesg, always on container-runtime 
address-space teardown,
hours to days before the panic:

```
BUG: Bad rss-counter state mm:ffff8934aa46a0c0 type:MM_ANONPAGES val:3 
Comm:containerd-shim Pid:3802600
BUG: non-zero pgtables_bytes on freeing mm: 8192
```

Also seen with `Comm:runc` (once with `non-zero pgtables_bytes: 49152`).

## Current state (2026-09-29)

| VM | Kernel | Uptime | Notes |
|---|---|---|---|
| vm-qat-0 | 7.0.0-1014 | 3 d 8 h | workload resident |
| vm-qat-1 | 7.0.0-1014 | 6 d 4 h | |
| vm-qat-2 | 6.17.0-1022 | 3 w 2 d | |
| vm-qat-3 | 6.17.0-1022 | 2 w 5 d | |
| vm-qat-4 | 7.0.0-1014 | 23 h | |
| prd-0..4 | 6.17.0-1018 / -1022 / 7.0.0-1014 | 4–77 d | |
| latest-0..2 | 6.14.0-1017 | ~275 d | control, zero panics |

No conclusion should be drawn yet about 7.0.0-1014. Historical intervals 
between panics on
affected kernels ranged **0.4 – 60.5 days** (median ~10), so current uptimes 
sit well inside
the range that affected kernels also survived.

## What we ruled out

- **Not slab corruption.** `slub_debug=FZPU` ran for 17 days with redzone + 
poison + user
  tracking active on all 545 caches (`red_zone=1 on 545 of 545`), across a 
panic, and
  produced **zero** reports. Consistent with the corrupted object being a 
vmalloc'd kernel
  stack rather than slab memory.
- **Not resource exhaustion.** At panic: ~1 GB of 15 GB used, CPU ~5%, OS disk 
IOPS 0%.
- **Not the hypervisor.** Azure Resource Health shows no platform event 
coinciding with any
  of the 62 panics. (Separately, vm-qat-4 hit an unrelated Azure 
host-degradation incident on
  2026-09-25 — a storage I/O hang with no panic. Different failure mode, 
excluded from the
  counts.)
- **Not hardware variation.** Every affected host and the control node 
vm-latest-0 are the
  same SKU (Standard_E2as_v5) on the same CPU (AMD EPYC 7763).
- **Userspace packages are implausible but not formally excluded** — see Limits.

## Limits of this evidence

Stated explicitly so nothing here is over-read:

- **The 26 Dec reboot changed two things, not one.** A manual `apt upgrade` ran 
on each host
  minutes before its reboot, carrying dmidecode, libpolkit-agent-1-0, dmeventd,
  initramfs-tools-core, netplan-generator, libsmartcols1, ssh-import-id and six 
python3
  packages, plus a 24.04.x point release of base-files. So kernel and userspace 
changed
  together. We consider userspace implausible as a cause — none of those 
packages ships
  kernel code, and these hosts boot initrdless via `GRUB_FORCE_PARTUUID`, making
  initramfs-tools inert — but the before/after boundary alone does not isolate 
the kernel.
- **Uptimes are lower bounds.** "Ran N days" is derived from the last 
timestamped line in
  each boot's console output, and "clean" means no panic was printed — a silent 
hang would
  be indistinguishable.
- **The maple-tree frames are stale stack remnants**, recovered from freed 
memory, not a live
  call trace. They establish what that stack had been doing, not what corrupted 
it.
- **The trigger correlation is across three environments, not a controlled 
test.** Query rate
  is a proxy for mmap/munmap rate, which we did not measure directly. And it is 
not
  proportional to panic frequency (see Trigger).
- **6.11.x is untested** on any affected host — it was installed but never 
booted, so the
  bisect window cannot currently be narrowed below 6.8 → 6.14.

A planned change will test the first point directly: moving these hosts to the
`linux-azure-6.8` GA track runs **current** userspace against the older kernel 
series. If the
panics stop, the kernel is isolated cleanly. QAT moves first, soaks two weeks, 
then prod.

## System details

- Ubuntu 24.04.3 LTS (noble)
- Azure `Standard_E2as_v5` — 2 vCPU, 16 GB, AMD EPYC 7763
- `CONFIG_VMAP_STACK=y`, `CONFIG_SCHED_STACK_END_CHECK=y`,
  `CONFIG_DEBUG_OBJECTS` **not set**, `CONFIG_KASAN` **not set**
- Kernel cmdline: `console=tty1 console=ttyS0 earlyprintk=ttyS0 panic=-1` + 
`crashkernel=`

Note: neither `CONFIG_DEBUG_OBJECTS` (which would catch a bad/on-stack 
`rcu_head` directly)
nor `CONFIG_KASAN` is available in the shipped Azure kernel, so we cannot 
instrument this
further without a custom build. **A test kernel with `DEBUG_OBJECTS_RCU_HEAD` 
would very
likely identify the offending `call_rcu()` caller immediately, and we can run 
one** on a host
that reproduces.

## Attachments

- `dmesg-20260926.txt` — full kernel ring buffer at panic, extracted by 
makedumpfile
- `dump.202609260602` — 1.7 GB vmcore (kdump, `-c -d 31`), available on request

kdump has since been reconfigured to `-c -d 14`, so the next capture will 
retain free and
zero pages — the memory a freed-kernel-stack bug most needs, and which `-d 31` 
stripped from
the dump above. We can supply a richer core after the next panic.
--- 
ProblemType: Bug
ApportVersion: 2.28.1-0ubuntu3.8
Architecture: amd64
AzureImageoffer: 0001-com-ubuntu-server-focal
AzureImagepublisher: canonical
AzureImagesku: 20_04-lts-gen2
AzureImageversion: 20.04.202308310
AzureVmsize: Standard_E2as_v5
CasperMD5CheckResult: unknown
CloudArchitecture: x86_64
CloudBuildName: server
CloudID: azure
CloudName: azure
CloudPlatform: azure
CloudRegion: eastus
CloudSerial: 20230831
CloudSubPlatform: seed-dir (/var/lib/waagent)
DistroRelease: Ubuntu 24.04
Package: linux-azure 7.0.0-1014.14~24.04.1
PackageArchitecture: amd64
ProcEnviron:
 LANG=C.UTF-8
 PATH=(custom, no user)
 SHELL=/bin/bash
 TERM=xterm-256color
 XDG_RUNTIME_DIR=<set>
ProcVersionSignature: User Name 7.0.0-1014.14~24.04.1-azure 7.0.14
Tags: cloud-image noble
Uname: Linux 7.0.0-1014-azure x86_64
UpgradeStatus: Upgraded to noble on 2024-09-20 (739 days ago)
UserGroups: adm audio cdrom dialout dip docker floppy lxd netdev plugdev sudo 
video
_MarkForUpload: True

** Affects: linux-azure (Ubuntu)
     Importance: Undecided
         Status: New


** Tags: amd64 apport-collected cloud-image kernel-bug noble regression-update

** Attachment added: "Full kernel ring buffer at panic, extracted by 
makedumpfile from the kdump vmcore"
   
https://bugs.launchpad.net/bugs/2168861/+attachment/6003603/+files/dmesg-20260926.txt

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2168861

Title:
  linux-azure: kernel panic in rcu_do_batch (NULL rcu_head on freed VMAP
  stack) — regression between 6.8 and 6.14

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-azure/+bug/2168861/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to