QEMU-Devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Peter Xu <peterx@redhat.com>
To: Yanfei Xu <isyanfei.xu@gmail.com>
Cc: Yanfei Xu <yanfei.xu@bytedance.com>,
	qemu-devel@nongnu.org, farosas@suse.de, lizhijian@fujitsu.com,
	Jinpu Wang <jinpu.wang@cloud.ionos.com>
Subject: Re: [PATCH 3/3] migration/rdma: Retry control sends on full queue
Date: Thu, 20 Aug 2026 09:47:53 -0400	[thread overview]
Message-ID: <aocFifqrtiv900Yr@x1.local> (raw)
In-Reply-To: <80b1c870-308f-4de4-ad53-c9ecfc91da24@gmail.com>

On Thu, Aug 20, 2026 at 06:27:12PM +0800, Yanfei Xu wrote:
> Hi peter,
> 
> On 2026/8/20 03:18, Peter Xu wrote:
> > On Mon, Aug 17, 2026 at 06:51:17PM +0800, Yanfei Xu wrote:
> > > RAM writes and control messages share the send queue.  If
> > > outstanding writes fill it, RDMA writes drain a completion and retry,
> > > but control sends fail the migration.
> > > 
> > > Drain one outstanding write and retry the control send on ENOMEM.
> > > 
> > > Signed-off-by: Yanfei Xu <yanfei.xu@bytedance.com>
> > > ---
> > >   migration/rdma.c | 12 +++++++++++-
> > >   1 file changed, 11 insertions(+), 1 deletion(-)
> > > 
> > > diff --git a/migration/rdma.c b/migration/rdma.c
> > > index 6e8436ccc1..d1f44a5f55 100644
> > > --- a/migration/rdma.c
> > > +++ b/migration/rdma.c
> > > @@ -1572,9 +1572,19 @@ static int qemu_rdma_post_send_control(RDMAContext *rdma, uint8_t *buf,
> > >           memcpy(wr->control + sizeof(RDMAControlHeader), buf, head->len);
> > >       }
> > > -
> > > +retry:
> > >       ret = ibv_post_send(rdma->qp, &send_wr, &bad_wr);
> > > +    if (ret == ENOMEM && rdma->nb_sent) {
> > > +        ret = qemu_rdma_block_for_wrid(rdma, RDMA_WRID_RDMA_WRITE, NULL);
> > > +        if (ret < 0) {
> > > +            error_setg(errp, "rdma migration: failed to make room for "
> > > +                       "control send");
> > > +            return -1;
> > > +        }
> > > +        goto retry;
> > > +    }
> > Looks also correct, but two questions:
> > 
> > - Should we provide a helper instead of duplicating the WRITE op handling?
> >    I believe only WRITE wrids can be on the fly.
> 
> Yes, control message is syncronized, only WRITE can be on the fly. A helper
> is a good suggestion. Will do.
> 
> > 
> > - Could ENOMEM be returned when nb_sent==0?  If that check applies to WRITE
> >    path too?
> 
> ENOMEM can be regarded as SQ is full to RDMA usage in qemu. Actually ENOMEM
> is determined by provider and could have other meaning like
> inline_data > qp->max_inline_data in mlx5. Based on qemu codes, I think it's
> fine without nb_sent==0

I'm not familiar with mlx5 impl that you're discussing here, but IIUC the
point is we should be able to capture all recoverable faults and retry,
meanwhile we should fail immediately on non-recoverable faults.

From the name of the errno (ENOMEM), I expect non-recoverable faults can
happen with it.. unless this is something special to libibverbs to
explicitly imply "queue full"..

I wonder if it means this nb_sent!=0 check should indeed make sense, but I
also wonder if we should add a number of retry so as to capture real ENOMEM
errors otherwise that is not recoverable?  As long as it won't keep
spinning with the same error then we should be good.

> 
> In addition, I encountered this bug when I attempt to send dirty pages
> belongs
> to same chunk in parallel. That could more efficiently utilizes throughput
> when
> many scattered page in one chunk, and quickly exhausts SQ's WRs. Will post a
> RFC with more data.

Sure, I hope that still makes sure different versions of a same page will
be still ordered.

Thanks,

-- 
Peter Xu



      reply	other threads:[~2026-08-20 13:48 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-17 10:51 [PATCH 0/3] migration/rdma: Misc RDMA migration fixes Yanfei Xu
2026-08-17 10:51 ` [PATCH 1/3] migration/rdma: Fix write-side shutdown Yanfei Xu
2026-08-19 14:58   ` Fabiano Rosas
2026-08-19 19:19   ` Peter Xu
2026-08-17 10:51 ` [PATCH 2/3] migration/rdma: Post initial receive before accepting Yanfei Xu
2026-08-19 19:16   ` Peter Xu
2026-08-20  5:46     ` Jinpu Wang
2026-08-20  6:22     ` Yanfei Xu
2026-08-20 13:37       ` Peter Xu
2026-08-17 10:51 ` [PATCH 3/3] migration/rdma: Retry control sends on full queue Yanfei Xu
2026-08-19 19:18   ` Peter Xu
2026-08-20 10:27     ` Yanfei Xu
2026-08-20 13:47       ` Peter Xu [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aocFifqrtiv900Yr@x1.local \
    --to=peterx@redhat.com \
    --cc=farosas@suse.de \
    --cc=isyanfei.xu@gmail.com \
    --cc=jinpu.wang@cloud.ionos.com \
    --cc=lizhijian@fujitsu.com \
    --cc=qemu-devel@nongnu.org \
    --cc=yanfei.xu@bytedance.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox