All of lore.kernel.org
 help / color / mirror / Atom feed
From: Leon Romanovsky <leon@kernel.org>
To: Serhat Kumral <serhatkumral1@gmail.com>
Cc: Bart Van Assche <bvanassche@acm.org>,
	Jason Gunthorpe <jgg@ziepe.ca>, Sagi Grimberg <sagi@grimberg.me>,
	linux-rdma@vger.kernel.org, target-devel@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v2] RDMA/srpt: Clamp the CQ size request to max_cqe
Date: Sun, 6 Sep 2026 12:30:30 +0300	[thread overview]
Message-ID: <20260906093030.GD13683@unreal> (raw)
In-Reply-To: <20260904200240.48976-1-serhatkumral1@gmail.com>

On Fri, Sep 04, 2026 at 11:02:40PM +0300, Serhat Kumral wrote:
> srpt_create_ch_ib() asks ib_cq_pool_get() for ch->rq_size + sq_size
> completion queue entries. rq_size is bounded by max_qp_wr, but sq_size
> comes from the per-port srp_sq_size configfs attribute, which is only
> validated against MAX_SRPT_SRQ_SIZE (65535) and never compared with
> dev->attrs.max_cqe. The largest configuration srpt accepts therefore
> asks for 128 + 65535 = 65663 entries, while rxe caps max_cqe at 32767
> and cxgb4 and ionic are in the same range.
> 
> Such a request cannot be met, and ib_cq_pool_get() does not reject it.
> It clamps every CQ it creates to max_cqe, so no CQ it adds to the pool
> can ever fit the request, and it keeps allocating batches until the
> allocation fails. A single SRP login against such a target exhausts
> memory:
> 
>   Out of memory and no killable processes...
>   Kernel panic - not syncing: System is deadlocked on memory
>   Workqueue: ib_cm cm_work_handler
>   Call Trace:
>    __vmalloc_node_range_noprof
>    vmalloc_user_noprof
>    rxe_queue_init
>    rxe_cq_from_init
>    rxe_create_cq
>    __ib_alloc_cq
>    ib_cq_pool_get
>    srpt_cm_req_recv.cold
> 
> Before commit c804af2c1d31 ("IB/srpt: use new shared CQ mechanism") the
> same request went to ib_alloc_cq_any(), which rejected it with -EINVAL.
> 
> The existing backoff, which halves sq_size when queue pair creation
> fails, only runs after ib_cq_pool_get() has returned, so shrink sq_size
> before asking for the CQ. rq_size never exceeds MAX_SRPT_RQ_SIZE (128),
> far below the max_cqe any device reports in practice, so the
> subtraction does not underflow. With the clamp the same login proceeds
> exactly like a correctly sized target.
> 
> Fixes: c804af2c1d31 ("IB/srpt: use new shared CQ mechanism")
> Signed-off-by: Serhat Kumral <serhatkumral1@gmail.com>
> ---
> Changes since v1:
> - Move the fix from the RDMA core to ib_srpt, as requested by Leon
>   Romanovsky.
> - Clamp sq_size before the CQ request instead of rejecting oversized
>   requests in ib_cq_pool_get().
> 
> v1: https://lore.kernel.org/linux-rdma/20260831171354.72140-1-serhatkumral1@gmail.com/
> 
>  drivers/infiniband/ulp/srpt/ib_srpt.c | 7 +++++++
>  1 file changed, 7 insertions(+)
> 
> diff --git a/drivers/infiniband/ulp/srpt/ib_srpt.c b/drivers/infiniband/ulp/srpt/ib_srpt.c
> index 7197d95f2216..58064a4c606b 100644
> --- a/drivers/infiniband/ulp/srpt/ib_srpt.c
> +++ b/drivers/infiniband/ulp/srpt/ib_srpt.c
> @@ -1867,6 +1867,13 @@ static int srpt_create_ch_ib(struct srpt_rdma_ch *ch)
>  	if (!qp_init)
>  		goto out;
>  
> +	/* The send and receive queues share a single CQ. */
> +	if (ch->rq_size + sq_size > attrs->max_cqe) {
> +		sq_size = attrs->max_cqe - ch->rq_size;

I don't think that Sashiko concern is real
https://sashiko.dev/#/patchset/20260904200240.48976-1-serhatkumral1@gmail.com
but it is worth to write it in a way to avoid triggering this warning.

For example
/* Catch drivers that incorrectly set ch->rq_size */
WARN_ON_ONCE(attrs->max_cqe < ch->rq_size)

Thanks

> +		pr_debug("reduced sq_size to %u because max_cqe is %u\n",
> +			 sq_size, attrs->max_cqe);
> +	}
> +
>  retry:
>  	ch->cq = ib_cq_pool_get(sdev->device, ch->rq_size + sq_size, -1,
>  				 IB_POLL_WORKQUEUE);
> 
> base-commit: cee9395acd8043be0644b25c34bfa86623f2b935
> -- 
> 2.53.0
> 

      reply	other threads:[~2026-09-06  9:30 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 20:02 [PATCH v2] RDMA/srpt: Clamp the CQ size request to max_cqe Serhat Kumral
2026-09-06  9:30 ` Leon Romanovsky [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260906093030.GD13683@unreal \
    --to=leon@kernel.org \
    --cc=bvanassche@acm.org \
    --cc=jgg@ziepe.ca \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sagi@grimberg.me \
    --cc=serhatkumral1@gmail.com \
    --cc=target-devel@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.