* [PATCH 0/2] migration/rdma: send small control messages inline @ 2026-09-01 15:51 Jack Wang 2026-09-01 15:51 ` [PATCH 1/2] " Jack Wang 2026-09-01 15:51 ` [PATCH 2/2] migration/rdma: avoid memcpy for inline control sends Jack Wang 0 siblings, 2 replies; 5+ messages in thread From: Jack Wang @ 2026-09-01 15:51 UTC (permalink / raw) To: qemu-devel; +Cc: Peter Xu, Fabiano Rosas, Li Zhijian, yanfei.xu, Jack Wang From: Jack Wang <jinpu.wang@cloud.ionos.com> Control-channel sends (registration requests, ram block replies, etc.) carry only a few dozen bytes, but qemu_rdma_post_send_control() still always memcpy()s them into a pre-registered scratch buffer and lets the HCA fetch them with a separate local memory read before sending. This series asks the QP for a little inline send space at creation time (patch 1) and then, once that plumbing exists, drops the mandatory memcpy() by pointing the SGEs directly at the header/payload for any message that fits inline (patch 2). Messages too large to inline keep using the old copy-into-registered-buffer path, so there is no wire protocol change. In testing, this doesn't show an obvious improvement in total migration time or downtime -- the control channel isn't the bottleneck there. Still, it seems like the right thing to do: inline sends are supported by modern HCAs and reduce per-message latency on the control channel by skipping the extra local memory read, so it should help control-plane responsiveness (e.g. registration round trips) even where it doesn't move the overall migration numbers. Jack Wang (2): migration/rdma: send small control messages inline migration/rdma: avoid memcpy for inline control sends migration/rdma.c | 82 +++++++++++++++++++++++++++++++++++++----------- 1 file changed, 63 insertions(+), 19 deletions(-) -- 2.43.0 ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 1/2] migration/rdma: send small control messages inline 2026-09-01 15:51 [PATCH 0/2] migration/rdma: send small control messages inline Jack Wang @ 2026-09-01 15:51 ` Jack Wang 2026-09-05 6:11 ` Yanfei Xu 2026-09-01 15:51 ` [PATCH 2/2] migration/rdma: avoid memcpy for inline control sends Jack Wang 1 sibling, 1 reply; 5+ messages in thread From: Jack Wang @ 2026-09-01 15:51 UTC (permalink / raw) To: qemu-devel; +Cc: Peter Xu, Fabiano Rosas, Li Zhijian, yanfei.xu, Jack Wang From: Jack Wang <jinpu.wang@cloud.ionos.com> Control-channel sends (registration requests, ram block replies, etc.) are tiny, but the HCA still does a separate memory read to fetch them before sending. Ask the QP for a little inline send space at creation time, and use it whenever a message is small enough to fit, so the HCA can just copy the bytes straight out of the work request instead. Falls back to the old path if the provider grants less inline room than we asked for, or none at all. No wire protocol change. Signed-off-by: Jack Wang <jinpu.wang@cloud.ionos.com> --- migration/rdma.c | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/migration/rdma.c b/migration/rdma.c index e976739fad3c..08f3b901be4a 100644 --- a/migration/rdma.c +++ b/migration/rdma.c @@ -66,6 +66,14 @@ static inline uint64_t rdma_merge_max(void) #define RDMA_CONTROL_MAX_BUFFER (512 * 1024) #define RDMA_CONTROL_MAX_COMMANDS_PER_MESSAGE 4096 +/* + * Requested max_inline_data for the QP: enough for a control header + * plus the largest fixed-size control payload, so small control + * messages can be sent inline instead of via a separate HCA-side + * memory read. + */ +#define RDMA_CONTROL_MAX_INLINE_DATA 512 + #define RDMA_CONTROL_VERSION_CURRENT 1 /* * Capabilities for negotiation. @@ -329,6 +337,7 @@ typedef struct RDMAContext { struct ibv_context *verbs; struct rdma_event_channel *channel; struct ibv_qp *qp; /* queue pair */ + uint32_t max_inline_data; /* max size for inline sends */ struct ibv_comp_channel *recv_comp_channel; /* recv completion channel */ struct ibv_comp_channel *send_comp_channel; /* send completion channel */ struct ibv_pd *pd; /* protection domain */ @@ -936,6 +945,14 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) attr.cap.max_recv_wr = 3; attr.cap.max_send_sge = 1; attr.cap.max_recv_sge = 1; + /* + * Ask for enough inline data to cover a control header plus the + * largest fixed-size control payload (RDMARegister/RDMACompress), + * so those sends can skip a local memory read on the HCA. The + * provider may grant less (or none); qemu_rdma_post_send_control() + * checks the actual granted size before using IBV_SEND_INLINE. + */ + attr.cap.max_inline_data = RDMA_CONTROL_MAX_INLINE_DATA; attr.send_cq = rdma->send_cq; attr.recv_cq = rdma->recv_cq; attr.qp_type = IBV_QPT_RC; @@ -945,6 +962,7 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) } rdma->qp = rdma->cm_id->qp; + rdma->max_inline_data = attr.cap.max_inline_data; return 0; } @@ -1447,6 +1465,10 @@ static int qemu_rdma_post_send_control(RDMAContext *rdma, uint8_t *buf, .num_sge = 1, }; + if (sge.length <= rdma->max_inline_data) { + send_wr.send_flags |= IBV_SEND_INLINE; + } + trace_rdma_post_send_control(control_desc(head->type)); /* -- 2.43.0 ^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH 1/2] migration/rdma: send small control messages inline 2026-09-01 15:51 ` [PATCH 1/2] " Jack Wang @ 2026-09-05 6:11 ` Yanfei Xu 2026-09-07 7:52 ` Jinpu Wang 0 siblings, 1 reply; 5+ messages in thread From: Yanfei Xu @ 2026-09-05 6:11 UTC (permalink / raw) To: Jack Wang, qemu-devel Cc: Peter Xu, Fabiano Rosas, Li Zhijian, yanfei.xu, Jack Wang On 2026/9/1 23:51, Jack Wang wrote: > From: Jack Wang <jinpu.wang@cloud.ionos.com> > > Control-channel sends (registration requests, ram block replies, etc.) > are tiny, but the HCA still does a separate memory read to fetch them > before sending. Ask the QP for a little inline send space at creation > time, and use it whenever a message is small enough to fit, so the > HCA can just copy the bytes straight out of the work request instead. > > Falls back to the old path if the provider grants less inline room > than we asked for, or none at all. No wire protocol change. > > Signed-off-by: Jack Wang <jinpu.wang@cloud.ionos.com> > --- > migration/rdma.c | 22 ++++++++++++++++++++++ > 1 file changed, 22 insertions(+) > > diff --git a/migration/rdma.c b/migration/rdma.c > index e976739fad3c..08f3b901be4a 100644 > --- a/migration/rdma.c > +++ b/migration/rdma.c > @@ -66,6 +66,14 @@ static inline uint64_t rdma_merge_max(void) > #define RDMA_CONTROL_MAX_BUFFER (512 * 1024) > #define RDMA_CONTROL_MAX_COMMANDS_PER_MESSAGE 4096 > > +/* > + * Requested max_inline_data for the QP: enough for a control header > + * plus the largest fixed-size control payload, so small control > + * messages can be sent inline instead of via a separate HCA-side > + * memory read. > + */ > +#define RDMA_CONTROL_MAX_INLINE_DATA 512 > + > #define RDMA_CONTROL_VERSION_CURRENT 1 > /* > * Capabilities for negotiation. > @@ -329,6 +337,7 @@ typedef struct RDMAContext { > struct ibv_context *verbs; > struct rdma_event_channel *channel; > struct ibv_qp *qp; /* queue pair */ > + uint32_t max_inline_data; /* max size for inline sends */ > struct ibv_comp_channel *recv_comp_channel; /* recv completion channel */ > struct ibv_comp_channel *send_comp_channel; /* send completion channel */ > struct ibv_pd *pd; /* protection domain */ > @@ -936,6 +945,14 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) > attr.cap.max_recv_wr = 3; > attr.cap.max_send_sge = 1; > attr.cap.max_recv_sge = 1; > + /* > + * Ask for enough inline data to cover a control header plus the > + * largest fixed-size control payload (RDMARegister/RDMACompress), > + * so those sends can skip a local memory read on the HCA. The > + * provider may grant less (or none); qemu_rdma_post_send_control() Seems it's actually opposite? From my understanding, QP allocation will failed if the required size is greater than provider's max inline cap. Then just sharing some my findings after learning about the inline feature: I found that requesting 512 bytes of inline data isn't free. Providers generally size the SQ for the worst-case WQE, so allowing a 512-byte inline payload may significantly increase both the maximum WQE size and the total SQ memory footprint. Regards, Yanfei > + * checks the actual granted size before using IBV_SEND_INLINE. > + */ > + attr.cap.max_inline_data = RDMA_CONTROL_MAX_INLINE_DATA; > attr.send_cq = rdma->send_cq; > attr.recv_cq = rdma->recv_cq; > attr.qp_type = IBV_QPT_RC; > @@ -945,6 +962,7 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) > } > > rdma->qp = rdma->cm_id->qp; > + rdma->max_inline_data = attr.cap.max_inline_data; > return 0; > } > > @@ -1447,6 +1465,10 @@ static int qemu_rdma_post_send_control(RDMAContext *rdma, uint8_t *buf, > .num_sge = 1, > }; > > + if (sge.length <= rdma->max_inline_data) { > + send_wr.send_flags |= IBV_SEND_INLINE; > + } > + > trace_rdma_post_send_control(control_desc(head->type)); > > /* ^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH 1/2] migration/rdma: send small control messages inline 2026-09-05 6:11 ` Yanfei Xu @ 2026-09-07 7:52 ` Jinpu Wang 0 siblings, 0 replies; 5+ messages in thread From: Jinpu Wang @ 2026-09-07 7:52 UTC (permalink / raw) To: Yanfei Xu; +Cc: qemu-devel, Peter Xu, Fabiano Rosas, Li Zhijian, yanfei.xu Hi Yanfei, You're right on both points, and I traced this through rdma-core to confirm. On the QP-allocation failure: in the mlx5 provider, mlx5_calc_send_wqe() computes the WQE size from the requested max_inline_data and returns -EINVAL if it exceeds the device's max SQ descriptor size. I also checked whether we could query the device's max inline capability up front instead of retrying blind — we can't.That's because the limit depends on the whole combination of requested QP caps, not a fixed device constant — so there's nothing to query in advance. On the SQ footprint point: agreed, it's a real per-QP memory cost regardless of whether 512B is ever used. Given all that, the fix has to be create-and-retry: request our size, and on failure retry with a reduced value (or 0) until ibv_create_qp() succeeds. I'll send a patch for that, along with shrinking the initial request from 512 down to something closer to our real payload sizes (max 40B). Regards, Jack On Sat, Sep 5, 2026 at 8:11 AM Yanfei Xu <isyanfei.xu@gmail.com> wrote: > > > On 2026/9/1 23:51, Jack Wang wrote: > > From: Jack Wang <jinpu.wang@cloud.ionos.com> > > > > Control-channel sends (registration requests, ram block replies, etc.) > > are tiny, but the HCA still does a separate memory read to fetch them > > before sending. Ask the QP for a little inline send space at creation > > time, and use it whenever a message is small enough to fit, so the > > HCA can just copy the bytes straight out of the work request instead. > > > > Falls back to the old path if the provider grants less inline room > > than we asked for, or none at all. No wire protocol change. > > > > Signed-off-by: Jack Wang <jinpu.wang@cloud.ionos.com> > > --- > > migration/rdma.c | 22 ++++++++++++++++++++++ > > 1 file changed, 22 insertions(+) > > > > diff --git a/migration/rdma.c b/migration/rdma.c > > index e976739fad3c..08f3b901be4a 100644 > > --- a/migration/rdma.c > > +++ b/migration/rdma.c > > @@ -66,6 +66,14 @@ static inline uint64_t rdma_merge_max(void) > > #define RDMA_CONTROL_MAX_BUFFER (512 * 1024) > > #define RDMA_CONTROL_MAX_COMMANDS_PER_MESSAGE 4096 > > > > +/* > > + * Requested max_inline_data for the QP: enough for a control header > > + * plus the largest fixed-size control payload, so small control > > + * messages can be sent inline instead of via a separate HCA-side > > + * memory read. > > + */ > > +#define RDMA_CONTROL_MAX_INLINE_DATA 512 > > + > > #define RDMA_CONTROL_VERSION_CURRENT 1 > > /* > > * Capabilities for negotiation. > > @@ -329,6 +337,7 @@ typedef struct RDMAContext { > > struct ibv_context *verbs; > > struct rdma_event_channel *channel; > > struct ibv_qp *qp; /* queue pair */ > > + uint32_t max_inline_data; /* max size for inline sends */ > > struct ibv_comp_channel *recv_comp_channel; /* recv completion channel */ > > struct ibv_comp_channel *send_comp_channel; /* send completion channel */ > > struct ibv_pd *pd; /* protection domain */ > > @@ -936,6 +945,14 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) > > attr.cap.max_recv_wr = 3; > > attr.cap.max_send_sge = 1; > > attr.cap.max_recv_sge = 1; > > + /* > > + * Ask for enough inline data to cover a control header plus the > > + * largest fixed-size control payload (RDMARegister/RDMACompress), > > + * so those sends can skip a local memory read on the HCA. The > > + * provider may grant less (or none); qemu_rdma_post_send_control() > > Seems it's actually opposite? From my understanding, QP allocation will > failed if > the required size is greater than provider's max inline cap. > > Then just sharing some my findings after learning about the inline > feature: I > found that requesting 512 bytes of inline data isn't free. Providers > generally > size the SQ for the worst-case WQE, so allowing a 512-byte inline > payload may > significantly increase both the maximum WQE size and the total SQ memory > footprint. > > Regards, > Yanfei > > > + * checks the actual granted size before using IBV_SEND_INLINE. > > + */ > > + attr.cap.max_inline_data = RDMA_CONTROL_MAX_INLINE_DATA; > > attr.send_cq = rdma->send_cq; > > attr.recv_cq = rdma->recv_cq; > > attr.qp_type = IBV_QPT_RC; > > @@ -945,6 +962,7 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) > > } > > > > rdma->qp = rdma->cm_id->qp; > > + rdma->max_inline_data = attr.cap.max_inline_data; > > return 0; > > } > > > > @@ -1447,6 +1465,10 @@ static int qemu_rdma_post_send_control(RDMAContext *rdma, uint8_t *buf, > > .num_sge = 1, > > }; > > > > + if (sge.length <= rdma->max_inline_data) { > > + send_wr.send_flags |= IBV_SEND_INLINE; > > + } > > + > > trace_rdma_post_send_control(control_desc(head->type)); > > > > /* ^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH 2/2] migration/rdma: avoid memcpy for inline control sends 2026-09-01 15:51 [PATCH 0/2] migration/rdma: send small control messages inline Jack Wang 2026-09-01 15:51 ` [PATCH 1/2] " Jack Wang @ 2026-09-01 15:51 ` Jack Wang 1 sibling, 0 replies; 5+ messages in thread From: Jack Wang @ 2026-09-01 15:51 UTC (permalink / raw) To: qemu-devel; +Cc: Peter Xu, Fabiano Rosas, Li Zhijian, yanfei.xu, Jack Wang From: Jack Wang <jinpu.wang@cloud.ionos.com> qemu_rdma_post_send_control() always copied the header and payload into a pre-registered scratch buffer before sending, even though inline sends don't need a registered region at all -- the HCA copies straight out of the given SGEs at post_send() time. When a message fits inline, point two SGEs directly at a stack-local header and the caller's payload instead of copying either into the scratch buffer. Confirmed against the mlx4/mlx5 driver source (providers/mlx5/qp.c:set_data_inl_seg(), providers/mlx4/qp.c) that inline SGEs never dereference lkey and are copied in a plain loop over num_sge, so this works for any SGE count the QP was created with. Falls back to the old copy-into-registered-buffer path when a message is too big to inline. This needs the QP to actually support 2 SGEs on a send, which the previous commit's QP creation didn't request (max_send_sge was still 1) -- bump it to 2. RDMA WRITEs are unaffected; they still always post exactly 1 SGE. Signed-off-by: Jack Wang <jinpu.wang@cloud.ionos.com> --- migration/rdma.c | 68 ++++++++++++++++++++++++++++++++---------------- 1 file changed, 45 insertions(+), 23 deletions(-) diff --git a/migration/rdma.c b/migration/rdma.c index 08f3b901be4a..88c4981804c3 100644 --- a/migration/rdma.c +++ b/migration/rdma.c @@ -943,7 +943,13 @@ static int qemu_rdma_alloc_qp(RDMAContext *rdma) attr.cap.max_send_wr = RDMA_SIGNALED_SEND_MAX; attr.cap.max_recv_wr = 3; - attr.cap.max_send_sge = 1; + /* + * RDMA WRITEs only ever use 1 SGE. Control sends use up to 2 when + * inlined (see qemu_rdma_post_send_control()): one for the header, + * one for the caller's payload, both pointing at unregistered + * memory that only inline sends can reference directly. + */ + attr.cap.max_send_sge = 2; attr.cap.max_recv_sge = 1; /* * Ask for enough inline data to cover a control header plus the @@ -1452,39 +1458,55 @@ static int qemu_rdma_post_send_control(RDMAContext *rdma, uint8_t *buf, int ret; RDMAWorkRequestData *wr = &rdma->wr_data[RDMA_WRID_CONTROL]; struct ibv_send_wr *bad_wr; - struct ibv_sge sge = { - .addr = (uintptr_t)(wr->control), - .length = head->len + sizeof(RDMAControlHeader), - .lkey = wr->control_mr->lkey, - }; + RDMAControlHeader net_head = *head; + uint32_t total_len = head->len + sizeof(RDMAControlHeader); + struct ibv_sge sge[2]; struct ibv_send_wr send_wr = { .wr_id = RDMA_WRID_SEND_CONTROL, .opcode = IBV_WR_SEND, .send_flags = IBV_SEND_SIGNALED, - .sg_list = &sge, + .sg_list = sge, .num_sge = 1, }; - if (sge.length <= rdma->max_inline_data) { - send_wr.send_flags |= IBV_SEND_INLINE; - } - trace_rdma_post_send_control(control_desc(head->type)); - /* - * We don't actually need to do a memcpy() in here if we used - * the "sge" properly, but since we're only sending control messages - * (not RAM in a performance-critical path), then its OK for now. - * - * The copy makes the RDMAControlHeader simpler to manipulate - * for the time being. - */ assert(head->len <= RDMA_CONTROL_MAX_BUFFER - sizeof(*head)); - memcpy(wr->control, head, sizeof(RDMAControlHeader)); - control_to_network((void *) wr->control); + control_to_network(&net_head); - if (buf) { - memcpy(wr->control + sizeof(RDMAControlHeader), buf, head->len); + if (total_len <= rdma->max_inline_data) { + /* + * Inline data is copied out of these SGEs by the HCA itself at + * post_send() time, so no registration (and no local copy into + * the pre-registered "control" buffer below) is needed -- point + * straight at the header on our stack and the caller's payload. + */ + sge[0].addr = (uintptr_t)&net_head; + sge[0].length = sizeof(net_head); + sge[0].lkey = 0; + + if (buf && head->len) { + sge[1].addr = (uintptr_t)buf; + sge[1].length = head->len; + sge[1].lkey = 0; + send_wr.num_sge = 2; + } + + send_wr.send_flags |= IBV_SEND_INLINE; + } else { + /* + * Too big to inline: the HCA will DMA-read this directly, which + * requires a registered region, so fall back to copying into + * the pre-registered "control" buffer. + */ + sge[0].addr = (uintptr_t)(wr->control); + sge[0].length = total_len; + sge[0].lkey = wr->control_mr->lkey; + + memcpy(wr->control, &net_head, sizeof(net_head)); + if (buf) { + memcpy(wr->control + sizeof(RDMAControlHeader), buf, head->len); + } } -- 2.43.0 ^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-07 7:53 UTC | newest] Thread overview: 5+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-01 15:51 [PATCH 0/2] migration/rdma: send small control messages inline Jack Wang 2026-09-01 15:51 ` [PATCH 1/2] " Jack Wang 2026-09-05 6:11 ` Yanfei Xu 2026-09-07 7:52 ` Jinpu Wang 2026-09-01 15:51 ` [PATCH 2/2] migration/rdma: avoid memcpy for inline control sends Jack Wang
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.