*** This bug is a duplicate of bug 2163303 ***
    https://bugs.launchpad.net/bugs/2163303

Public bug reported:

# linux-firmware 2.29: DMCUB firmware fails to load on Navi 21 (RX
6800), causing display corruption and GPU hangs in multi-monitor setups

## Summary

The `sienna_cichlid_dmcub.bin` shipped in `linux-firmware`
`20240318.git3b128b60-0ubuntu2.29` fails to load on Navi 21 (Radeon RX 6800,
DCN 3.0). The kernel reports an invalid DMUB version (`0x01000000`) and
`Wait for DMUB auto-load failed: 3`, then continues operating with the display
microcontroller in a degraded state.

With more than one display active this leads to display plane corruption,
`amdgpu_dm_commit_planes` warnings, MPC register wait timeouts, SMU
communication loss, GPU ring timeouts and eventually a full system freeze.

Downgrading **only** that one firmware file to the version from
`20240318.git3b128b60-0ubuntu2.26` fixes the problem completely.

## Package / versions

- Distro: Linux Mint 22.3 (Ubuntu 24.04 noble base)
- Package: `linux-firmware`
- Bad version: `20240318.git3b128b60-0ubuntu2.29`
- Good version: `20240318.git3b128b60-0ubuntu2.26`
- Affected file: `/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst`
  - `.29` md5: `8783824f37745ec5d53ee8a2d71b18f7`
  - `.26` md5: `041e8ee4f578b5eccb1bb89a2f14b1db`

Of the 12 `sienna_cichlid_*` firmware files, this is the **only** one that
differs between `.26` and `.29`. All others are byte-identical.

## Hardware

- GPU: `Advanced Micro Devices, Inc. [AMD/ATI] Navi 21 [Radeon RX 6800/6800 XT 
/ 6900 XT] [1002:73bf] (rev c3)`
  (Sapphire, 16 GB, DCN 3.0, Display Core v3.2.369)
- Motherboard: ASUS ROG STRIX B350-F GAMING, BIOS 5603
- Displays: 3 monitors
  - `DP-1` 1920x1080@60 (Philips)
  - `DP-2` 1920x1080@60
  - `HDMI-A-1` [email protected]
- Session: X11 (Cinnamon)

## Not kernel-dependent

The failure follows the firmware, not the kernel. All three installed kernels
reproduce it while the `.29` firmware is in place:

| Kernel | DMUB version reported | GPU errors in that boot |
|---|---|---|
| 6.14.0-37-generic | `0x01000000` | yes |
| 7.0.0-28-generic  | `0x01000000` | yes (47) |
| 7.0.0-29-generic  | `0x01000000` | yes (21 / 23 / 24 / 212) |
| 7.0.0-29-generic  | `0x02020020` (after firmware downgrade) | **0** |

## Steps to reproduce

1. Navi 21 GPU with `linux-firmware` `2.29` installed.
2. Connect 3 displays (2x DisplayPort, 1x HDMI) and boot into a graphical 
session.
3. Use the desktop normally. Anything that triggers a connector re-probe
   (for example running `xrandr --query`, or the desktop environment
   re-reading the display configuration) can trigger the failure.

## Expected

DMUB firmware loads and reports a valid version; displays work normally.

## Actual

### 1. DMUB fails to load at every boot

```
amdgpu 0000:0b:00.0: [drm] Loading DMUB firmware via PSP: version=0x01000000
amdgpu 0000:0b:00.0: [drm] Wait for DMUB auto-load failed: 3
amdgpu 0000:0b:00.0: [drm] DMUB hardware initialized: version=0x01000000
```

`0x01000000` is not a valid DMUB firmware version.

### 2. Visible symptom

Vertical coloured stripes covering roughly half of one display. Application
windows render *behind* the stripes, while the mouse cursor draws correctly
*in front* of them — consistent with the hardware cursor plane still being
updated while the main plane composition is stuck.

Never reproduced with a single display active (e.g. recovery mode).
Never reproduced under Windows on the same hardware and cables.

### 3. Kernel warning in the plane commit path

```
WARNING: drivers/gpu/drm/amd/amdgpu/../display/amdgpu_dm/amdgpu_dm.c:10181
  at amdgpu_dm_commit_planes+0x11c8/0x1740 [amdgpu]
WARNING: drivers/gpu/drm/amd/amdgpu/../display/amdgpu_dm/amdgpu_dm.c:9567
  at amdgpu_dm_commit_planes+0x11cf/0x1740 [amdgpu]

Call Trace:
 amdgpu_dm_commit_streams+0x658/0x8b0 [amdgpu]
 amdgpu_dm_atomic_commit_tail+0xb9/0xda0 [amdgpu]
 commit_tail+0xc9/0x1b0
 drm_atomic_helper_commit+0x132/0x160
 drm_atomic_commit+0xaf/0xf0
 dm_restore_drm_connector_state+0x101/0x1b0 [amdgpu]
 handle_hpd_irq_helper+0x24e/0x310 [amdgpu]
 handle_hpd_irq+0xe/0x20 [amdgpu]
 dm_irq_work_func+0x19/0x30 [amdgpu]
 process_one_work+0x1af/0x430
 worker_thread+0x1bf/0x350
```

### 4. MPC never goes idle

```
amdgpu 0000:0b:00.0: [drm] REG_WAIT timeout 1us * 100000 tries - 
mpc2_assert_idle_mpcc line:481
amdgpu 0000:0b:00.0: [drm] REG_WAIT timeout 1us * 10 tries - optc3_lock line:128
workqueue: dm_irq_work_func [amdgpu] hogged CPU for >10000us 11 times
[drm:dcn20_wait_for_blank_complete [amdgpu]] *ERROR* DC: failed to blank crtc!
amdgpu 0000:0b:00.0: [drm] *ERROR* [CRTC:470:crtc-0] flip_done timed out
amdgpu 0000:0b:00.0: [drm] *ERROR* [CRTC:470:crtc-0] commit wait timed out
```

### 5. Cascade into SMU loss and GPU hang

```
amdgpu 0000:0b:00.0: SMU: No response msg_reg: 22 resp_reg: 0     (x210 in one 
boot)
amdgpu 0000:0b:00.0: Failed to disable gfxoff!
amdgpu 0000:0b:00.0: Failed to set workload mask 0x00000001
amdgpu 0000:0b:00.0: (-62) failed to disable fullscreen 3D power profile mode
amdgpu 0000:0b:00.0: ring gfx_0.0.0 timeout, signaled seq=25441, emitted 
seq=25441
amdgpu 0000:0b:00.0: ring sdma2 timeout, signaled seq=620, emitted seq=622
amdgpu 0000:0b:00.0: [drm:amdgpu_ring_test_helper [amdgpu]] *ERROR* ring 
kiq_0.2.1.0 test failed (-110)
amdgpu 0000:0b:00.0: [drm] device wedged, but recovered through reset
```

### 6. PCIe AER errors from the GPU and its audio function

```
pcieport 0000:00:03.1: AER: Multiple Uncorrectable (Non-Fatal) error message 
received from 0000:0b:00.1
amdgpu 0000:0b:00.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal), 
type=Transaction Layer, (Requester ID)
amdgpu 0000:0b:00.0:   device [1002:73bf] error status/mask=00100000/00000000
snd_hda_intel 0000:0b:00.1: AER:   TLP Header: 0x40001001 0x0000000c 0xfca2000c 
0x00000000
pcieport 0000:0a:00.0: AER: device recovery failed
```

The system then freezes hard — screens go black, nothing further is written to
the journal, and a forced power-off is required.

## Workarounds that did NOT help

- Booting the previous kernel (`7.0.0-28-generic`) or `6.14.0-37-generic`
- Xorg `Option "EnablePageFlip" "off"` / `"TearFree" "false"` / 
`"AsyncFlipSecondaries" "false"`
- `amdgpu.ppfeaturemask` with GFXOFF disabled
- `amdgpu.dcdebugmask=0x10` + `amdgpu.runpm=0` + `amdgpu.gpu_recovery=1`

Note: `amdgpu.dc=0` is **not** a viable workaround on this hardware — Display
Core is mandatory for DCN, and the parameter results in no display output at 
all.

## Workaround that DOES work

Replace only `/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst` with the file
from `linux-firmware` `20240318.git3b128b60-0ubuntu2.26`, regenerate the
initramfs and reboot:

```
apt-get download linux-firmware=20240318.git3b128b60-0ubuntu2.26
dpkg-deb -x linux-firmware_*2.26_amd64.deb fw26/
cp fw26/lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst \
   /lib/firmware/amdgpu/sienna_cichlid_dmcub.bin.zst
update-initramfs -u -k all
```

After this the DMUB loads correctly and all errors disappear:

```
amdgpu 0000:0b:00.0: [drm] Loading DMUB firmware via PSP: version=0x02020020
amdgpu 0000:0b:00.0: [drm] DMUB hardware initialized: version=0x02020020
```

No `Wait for DMUB auto-load failed`, no plane commit warnings, no SMU timeouts,
no freezes, with all 3 displays active.

## Impact

Any Navi 21 (RX 6800 / 6800 XT / 6900 XT) system on Ubuntu 24.04 / derivatives
with `linux-firmware` `2.29` and more than one display is likely affected. The
symptom is a hard system freeze, so it can cause data loss.

## Note on the 2.27 version

`20240318.git3b128b60-0ubuntu2.27` (the version that introduced the update on
this system) is no longer available in the archive; only `.26`, `.29` and the
original `2` are. The regression was therefore bisected to the `.26` → `.29`
range rather than to an exact intermediate version. The `.27` and `.28`
changelog entries contain numerous "DMCUB updates for various AMDGPU ASICs"
commits.

** Affects: linux-firmware (Ubuntu)
     Importance: Undecided
         Status: New

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2164690

Title:
  linux-firmware 2.29: DMCUB firmware fails to load on Navi 21 (RX
  6800), causing display corruption and GPU hangs in multi-monitor
  setups

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-firmware/+bug/2164690/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to