On Tue, Aug 11, 2026 at 10:25 PM Liu, Changcheng <[email protected]> wrote: > > On Tue, Aug 11, 2026 at 10:37:28AM -0400, Stefan Hajnoczi wrote: > > External email: Use caution opening links or attachments > > > > > > On Wed, Aug 5, 2026 at 10:40 PM Liu, Changcheng <[email protected]> wrote: > > > > > > On Wed, Aug 05, 2026 at 01:45:58PM -0400, Stefan Hajnoczi wrote: > > > > > > > > On Sat, Aug 1, 2026 at 2:21 AM Liu, Changcheng <[email protected]> > > > > wrote: > > > > > > > > > > When creating many virtio-blk devices, probe starts failing with > > > > > -ENOSPC (-28) because the system runs out of interrupt vectors: > > > > > > > > > > virtio_blk virtioNNN: probe with driver virtio_blk failed with > > > > > error -28 > > > > > > > > > > By default virtio-blk uses managed IRQ affinity, which reserves an > > > > > interrupt vector on every CPU for each device (about nr_cpus vectors > > > > > per > > > > > device). On a host with many CPUs and many devices this exhausts the > > > > > vectors long before all devices are probed. > > > > > > > > > > Add use_irq_affinity (default true, no behaviour change). Set it to 0 > > > > > to > > > > > use unmanaged interrupts, so each device only uses a couple of vectors > > > > > instead of one per CPU, allowing far more devices to probe. > > > > > > > > > > Signed-off-by: Liu, Changcheng <[email protected]> > > > > > > > > There is already a num_request_queues module parameter for cases where > > > > the user wishes to reduce the number of virtqueues. Did you benchmark > > > > that and decide the performance of many queues sharing a single irq > > > > makes it worth adding another module parameter? > > > > > > > > Stefan > > > > > > num_request_queues does not address this case because each virtio-blk > > > device > > > already has only one request virtqueue when the failure occurs, so the > > > queue > > > count cannot be reduced further. > > > > Okay. > > > > > The issue is reproduced after probing approximately 800 virtio-blk > > > devices on > > > a 64-core bare-metal host. With managed IRQ affinity, the IRQ-vector > > > reservations > > > are eventually exhausted and probing additional devices fails with > > > -ENOSPC. > > > Disabling managed IRQ affinity avoids these per-CPU vector reservations > > > and > > > allows more devices to be probed successfully. > > > > I don't follow. My understanding was that non-managed IRQ vectors are > > allocated by request_irq(), which is called during virtio_find_vqs(). > > Probing 800 devices would still require 800 * (1 virtqueue irq + 1 > > config change irq) = 1,600 vectors. How come non-managed IRQs do not > > hit the limit here? > > > > Thanks, > > Stefan > > > > With one request queue, virtio-blk allocates two MSI-X interrupts: an > unmanaged configuration interrupt and a managed request-queue interrupt. > > The request-queue interrupt's managed affinity mask covers all possible > CPUs. On x86, irq_matrix_reserve_managed() reserves one vector slot on > every CPU in that mask so the interrupt can migrate safely during CPU > hotplug. Consequently, every virtio-blk device consumes one managed > vector slot on every CPU for its lifetime. > > Once any CPU exhausts its roughly 200 usable vector slots, managed vector > allocation fails even if other CPUs still have capacity. virtio-pci then > falls back to its shared-vector MSI-X allocation, for which the affinity > descriptor is discarded and the interrupts are unmanaged. Allocation can > continue using capacity on the other CPUs, explaining the observed limit > of approximately 800 devices. > > When managed affinity is disabled, both interrupts are activated on > individual CPUs and distributed by the vector allocator. APIC vector > numbers are per-CPU, so the same vector number can be reused on different > CPUs. On a 64-CPU system, 800 devices therefore consume approximately > 800 * 2 / 64 = 25 vector slots per CPU.
Hi Changcheng, Thanks for the explanation. Can you update the commit message to mention that non-managed irqs stay below the limit (800 * 2 / 64 = 25 vector slots per CPU) because they are allocated to the CPU with the lowest number by the vector allocator? Other than that: Reviewed-by: Stefan Hajnoczi <[email protected]>
