Hi all,

We run 26.06 (plus a few backported fixes, details below) with linux-cp/FRR on 
four routers. For several months we've seen sporadic worker crashes, days to 
weeks apart. On 2026-10-05 we captured the first two cores with full symbols, 
from two routers within 12 minutes of each other. Both show a worker following 
a stale reference.

*Setup (both routers):* DPDK on Intel 10G (ixgbe/i40e), 2 workers, linux-cp 
with lcp-sync and lcp-auto-subint, FRR (OSPFv2/v3 and BGP full tables), sFlow 
on the transit ports. The inter-site links are an L2 bridge-domain per site 
pair (learn 1 flood 1 uu-flood 0), each with a BVI loopback (linux-cp mirrored) 
and VXLAN tunnels to the other sites.

*Crash 1 (router A), worker 1, SIGSEGV:*

vlib_next_frame_change_ownership (next_index=18141)  vlib/main.c:234
vlib_get_next_frame_internal                          vlib/main.c:308
enqueue_one / vlib_buffer_enqueue_to_next_fn          vlib/buffer_funcs.c:21
ip4_rewrite_inline                                    vnet/ip/ip4_forward.c:2515

The packet's adj_index[VLIB_TX] was 5121, and that adjacency was garbage: 
lookup_next_index=143, ia_nh_proto=4, rewrite_header = { sw_if_index=278152, 
next_index=18141, max_l3_packet_bytes=18143, data_bytes=4 }. It looks like a 
freed and reused pool slot. The packet was VXLAN (UDP 4789) from our loopback 
to the remote VTEP, i.e. traffic from the bridge-domain above. The main thread 
was in vlib_rpc_call_main_thread_process -> barrier_sync, and the other worker 
was parked in barrier_check.

*Crash 2 (router B), worker 0, SIGSEGV at address 0x20:*

eth_input_process_frame_dmac_check (ei=0x0, have_sec_dmac=0)   
vnet/ethernet/node.c:709
eth_input_process_frame (main_is_l3=1, dmac_check=1)          node.c:891
eth_input_single_int / ethernet_input_node_fn

The single buffer in the frame had already been freed: ref_count=0, 
current_data=14, current_length=65522 (i.e. -14), RX/TX sw_if_index={0,0}, 
packet data all zeros. So hi resolved to local0, which has no ethernet 
interface, giving ei=NULL. The main thread was idle in epoll_wait, and the 
other worker was in tap-input.

An earlier crash on router B (2026-09-30, no core) had the same signature as 
crash 1: the same function, offset and code bytes in ip4_rewrite.

In all cases the main thread wasn't modifying anything at the moment of the 
crash, so the free happened earlier and a stale index survived in a frame or 
buffer. Our current suspicion is the BVI + VXLAN bridge path (l2-flood with 
vlib_buffer_clone for BUM traffic), but we haven't proven it. We've seen the L2 
fixes merged since 26.06 (3294213d1, 2543d3fe5, 6ad56defd) and will test them, 
but none obviously explains a freed adjacency or a freed buffer.

*Build:* v26.06 + 1bba2c01, 7dc47fc1 (44230), our 46937, the ip6_nd hunk of 
46038, 45954, 45955, 46765, and a local LACP clock change (46968). None of 
these touch L2, VXLAN, ethernet-input or the buffer code.

We have both cores (14 GB and 20 GB) and the matching debug packages, and can 
run any gdb commands you'd like, or share specific memory. Any pointers on 
where to look, or debug options to enable, would be appreciated.

Regards,
Pieter Meyer
-=-=-=-=-=-=-=-=-=-=-=-
Links: You receive all messages sent to this group.
View/Reply Online (#27229): https://lists.fd.io/g/vpp-dev/message/27229
Mute This Topic: https://lists.fd.io/mt/121592811/21656
Group Owner: [email protected]
Unsubscribe: https://lists.fd.io/g/vpp-dev/leave/14379924/21656/631435203/xyzzy 
[[email protected]]
-=-=-=-=-=-=-=-=-=-=-=-

Reply via email to