Hi peter,
On 2026/8/20 03:18, Peter Xu wrote:
On Mon, Aug 17, 2026 at 06:51:17PM +0800, Yanfei Xu wrote:
RAM writes and control messages share the send queue. If
outstanding writes fill it, RDMA writes drain a completion and retry,
but control sends fail the migration.
Drain one outstanding write and retry the control send on ENOMEM.
Signed-off-by: Yanfei Xu <[email protected]>
---
migration/rdma.c | 12 +++++++++++-
1 file changed, 11 insertions(+), 1 deletion(-)
diff --git a/migration/rdma.c b/migration/rdma.c
index 6e8436ccc1..d1f44a5f55 100644
--- a/migration/rdma.c
+++ b/migration/rdma.c
@@ -1572,9 +1572,19 @@ static int qemu_rdma_post_send_control(RDMAContext
*rdma, uint8_t *buf,
memcpy(wr->control + sizeof(RDMAControlHeader), buf, head->len);
}
-
+retry:
ret = ibv_post_send(rdma->qp, &send_wr, &bad_wr);
+ if (ret == ENOMEM && rdma->nb_sent) {
+ ret = qemu_rdma_block_for_wrid(rdma, RDMA_WRID_RDMA_WRITE, NULL);
+ if (ret < 0) {
+ error_setg(errp, "rdma migration: failed to make room for "
+ "control send");
+ return -1;
+ }
+ goto retry;
+ }
Looks also correct, but two questions:
- Should we provide a helper instead of duplicating the WRITE op handling?
I believe only WRITE wrids can be on the fly.
Yes, control message is syncronized, only WRITE can be on the fly. A
helper is a good suggestion. Will do.
- Could ENOMEM be returned when nb_sent==0? If that check applies to WRITE
path too?
ENOMEM can be regarded as SQ is full to RDMA usage in qemu. Actually ENOMEM
is determined by provider and could have other meaning like
inline_data > qp->max_inline_data in mlx5. Based on qemu codes, I think it's
fine without nb_sent==0
In addition, I encountered this bug when I attempt to send dirty pages
belongs
to same chunk in parallel. That could more efficiently utilizes
throughput when
many scattered page in one chunk, and quickly exhausts SQ's WRs. Will post a
RFC with more data.
Regards,
Yanfei
Thanks,
+
if (ret > 0) {
error_setg(errp, "Failed to use post IB SEND for control");
return -1;
--
2.20.1