Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing
@ 2026-09-19 10:08 lirongqing
  2026-09-19 10:17 ` sashiko-bot
  2026-09-28 18:38 ` Leon Romanovsky
  0 siblings, 2 replies; 4+ messages in thread
From: lirongqing @ 2026-09-19 10:08 UTC (permalink / raw)
  To: Leon Romanovsky, Jason Gunthorpe, linux-rdma; +Cc: Li RongQing

From: Li RongQing <lirongqing@baidu.com>

In mlx5_poll_one, a MLX5_CQE_SIG_ERR completion looks up the signature
context in dev->sig_mrs via xa_load() but dereferences the result without
a NULL check. If the corresponding MR has been concurrently deregistered
(which removes the xarray entry without holding the CQ lock), xa_load()
returns NULL and the subsequent get_sig_err_item(sig_err_cqe, &sig->err_item)
crashes in atomic context under xa_lock.

Add a NULL check that drops the lock, logs a warning, and repolls.

Signed-off-by: Li RongQing <lirongqing@baidu.com>
---
 drivers/infiniband/hw/mlx5/cq.c | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/drivers/infiniband/hw/mlx5/cq.c b/drivers/infiniband/hw/mlx5/cq.c
index 49b4bf1..b0b1177 100644
--- a/drivers/infiniband/hw/mlx5/cq.c
+++ b/drivers/infiniband/hw/mlx5/cq.c
@@ -563,6 +563,13 @@ static int mlx5_poll_one(struct mlx5_ib_cq *cq,
 		xa_lock(&dev->sig_mrs);
 		sig = xa_load(&dev->sig_mrs,
 				mlx5_base_mkey(be32_to_cpu(sig_err_cqe->mkey)));
+		if (!sig) {
+			xa_unlock(&dev->sig_mrs);
+			mlx5_ib_warn(dev, "CQN: 0x%x Got SIGERR on unknown mkey: 0x%x\n",
+				     cq->mcq.cqn,
+				     be32_to_cpu(sig_err_cqe->mkey));
+			goto repoll;
+		}
 		get_sig_err_item(sig_err_cqe, &sig->err_item);
 		sig->sig_err_exists = true;
 		sig->sigerr_count++;
-- 
2.9.4


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing
  2026-09-19 10:08 [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing lirongqing
@ 2026-09-19 10:17 ` sashiko-bot
  2026-09-28 18:38 ` Leon Romanovsky
  1 sibling, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-09-19 10:17 UTC (permalink / raw)
  To: lirongqing; +Cc: linux-rdma

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] rdma/mlx5: un-ratelimited warning in CQ polling loop causes log flood
--

commit c105d16aafdc1a586e27b8a4c8ea42b5926b0203
Author: Li RongQing <lirongqing@baidu.com>

RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing

This commit adds a NULL check after looking up a signature context in
mlx5_poll_one() to prevent a crash when a memory region is concurrently
deregistered. If the lookup fails, it drops the lock, logs a warning, and
fetches the next completion.

> diff --git a/drivers/infiniband/hw/mlx5/cq.c b/drivers/infiniband/hw/mlx5/cq.c
> index 49b4bf148a4a0..b0b1177077aa0 100644
> --- a/drivers/infiniband/hw/mlx5/cq.c
> +++ b/drivers/infiniband/hw/mlx5/cq.c
> @@ -563,6 +563,13 @@ static int mlx5_poll_one(struct mlx5_ib_cq *cq,
>  		xa_lock(&dev->sig_mrs);
>  		sig = xa_load(&dev->sig_mrs,
>  				mlx5_base_mkey(be32_to_cpu(sig_err_cqe->mkey)));
> +		if (!sig) {
> +			xa_unlock(&dev->sig_mrs);
> +			mlx5_ib_warn(dev, "CQN: 0x%x Got SIGERR on unknown mkey: 0x%x\n",
> +				     cq->mcq.cqn,
> +				     be32_to_cpu(sig_err_cqe->mkey));
> +			goto repoll;

[Severity: Medium]
Could this un-ratelimited warning lead to console flooding and CPU soft
lockups?

A stream of MLX5_CQE_SIG_ERR completions with unknown mkeys (which can be
triggered by remote peers sending invalid data to a registered memory region)
causes mlx5_poll_one() to repeatedly call mlx5_ib_warn() and jump to the
repoll label.

Because goto repoll fetches the next completion queue entry without returning,
it bypasses the CQ polling budget entirely. This could allow a remote attacker
to intentionally stream invalid packets to flood the queue, weaponizing the
loop into a denial of service vector.

Should this use a ratelimited print, and possibly avoid bypassing the polling
budget?

> +		}
>  		get_sig_err_item(sig_err_cqe, &sig->err_item);
>  		sig->sig_err_exists = true;
>  		sig->sigerr_count++;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919100806.2320-1-lirongqing@baidu.com?part=1

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing
  2026-09-19 10:08 [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing lirongqing
  2026-09-19 10:17 ` sashiko-bot
@ 2026-09-28 18:38 ` Leon Romanovsky
  2026-09-29  1:21   ` 答复: [????] " Li,Rongqing
  1 sibling, 1 reply; 4+ messages in thread
From: Leon Romanovsky @ 2026-09-28 18:38 UTC (permalink / raw)
  To: lirongqing; +Cc: Jason Gunthorpe, linux-rdma

On Sat, Sep 19, 2026 at 06:08:06PM +0800, lirongqing wrote:
> From: Li RongQing <lirongqing@baidu.com>
> 
> In mlx5_poll_one, a MLX5_CQE_SIG_ERR completion looks up the signature
> context in dev->sig_mrs via xa_load() but dereferences the result without
> a NULL check. If the corresponding MR has been concurrently deregistered

Is this real scenario? Integrity MRs are in-kernel MRs, they shouldn't
be concurrently deregistered.

Thanks

> (which removes the xarray entry without holding the CQ lock), xa_load()
> returns NULL and the subsequent get_sig_err_item(sig_err_cqe, &sig->err_item)
> crashes in atomic context under xa_lock.
> 
> Add a NULL check that drops the lock, logs a warning, and repolls.
> 
> Signed-off-by: Li RongQing <lirongqing@baidu.com>
> ---
>  drivers/infiniband/hw/mlx5/cq.c | 7 +++++++
>  1 file changed, 7 insertions(+)
> 
> diff --git a/drivers/infiniband/hw/mlx5/cq.c b/drivers/infiniband/hw/mlx5/cq.c
> index 49b4bf1..b0b1177 100644
> --- a/drivers/infiniband/hw/mlx5/cq.c
> +++ b/drivers/infiniband/hw/mlx5/cq.c
> @@ -563,6 +563,13 @@ static int mlx5_poll_one(struct mlx5_ib_cq *cq,
>  		xa_lock(&dev->sig_mrs);
>  		sig = xa_load(&dev->sig_mrs,
>  				mlx5_base_mkey(be32_to_cpu(sig_err_cqe->mkey)));
> +		if (!sig) {
> +			xa_unlock(&dev->sig_mrs);
> +			mlx5_ib_warn(dev, "CQN: 0x%x Got SIGERR on unknown mkey: 0x%x\n",
> +				     cq->mcq.cqn,
> +				     be32_to_cpu(sig_err_cqe->mkey));
> +			goto repoll;
> +		}
>  		get_sig_err_item(sig_err_cqe, &sig->err_item);
>  		sig->sig_err_exists = true;
>  		sig->sigerr_count++;
> -- 
> 2.9.4
> 

^ permalink raw reply	[flat|nested] 4+ messages in thread

* 答复: [????] Re: [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing
  2026-09-28 18:38 ` Leon Romanovsky
@ 2026-09-29  1:21   ` Li,Rongqing
  0 siblings, 0 replies; 4+ messages in thread
From: Li,Rongqing @ 2026-09-29  1:21 UTC (permalink / raw)
  To: Leon Romanovsky; +Cc: Jason Gunthorpe, linux-rdma@vger.kernel.org


> On Sat, Sep 19, 2026 at 06:08:06PM +0800, lirongqing wrote:
> > From: Li RongQing <lirongqing@baidu.com>
> >
> > In mlx5_poll_one, a MLX5_CQE_SIG_ERR completion looks up the signature
> > context in dev->sig_mrs via xa_load() but dereferences the result
> > without a NULL check. If the corresponding MR has been concurrently
> > deregistered
> 
> Is this real scenario? Integrity MRs are in-kernel MRs, they shouldn't be
> concurrently deregistered.
> 
Thanks for the review. You're right ― the concurrent-deregister scenario I described is not reachable in normal use: MLX5_CQE_SIG_ERR is only generated for signature (IB_MR_TYPE_INTEGRITY) MRs, whose mkey is always present in dev->sig_mrs while the operation is in flight, and the null mkey is an ODP-only special case that is never inserted into sig_mrs.

Thanks

[Li,Rongqing] 


> Thanks
> 
> > (which removes the xarray entry without holding the CQ lock),
> > xa_load() returns NULL and the subsequent
> > get_sig_err_item(sig_err_cqe, &sig->err_item) crashes in atomic context under
> xa_lock.
> >
> > Add a NULL check that drops the lock, logs a warning, and repolls.
> >
> > Signed-off-by: Li RongQing <lirongqing@baidu.com>
> > ---
> >  drivers/infiniband/hw/mlx5/cq.c | 7 +++++++
> >  1 file changed, 7 insertions(+)
> >
> > diff --git a/drivers/infiniband/hw/mlx5/cq.c
> > b/drivers/infiniband/hw/mlx5/cq.c index 49b4bf1..b0b1177 100644
> > --- a/drivers/infiniband/hw/mlx5/cq.c
> > +++ b/drivers/infiniband/hw/mlx5/cq.c
> > @@ -563,6 +563,13 @@ static int mlx5_poll_one(struct mlx5_ib_cq *cq,
> >  		xa_lock(&dev->sig_mrs);
> >  		sig = xa_load(&dev->sig_mrs,
> >  				mlx5_base_mkey(be32_to_cpu(sig_err_cqe->mkey)));
> > +		if (!sig) {
> > +			xa_unlock(&dev->sig_mrs);
> > +			mlx5_ib_warn(dev, "CQN: 0x%x Got SIGERR on unknown mkey:
> 0x%x\n",
> > +				     cq->mcq.cqn,
> > +				     be32_to_cpu(sig_err_cqe->mkey));
> > +			goto repoll;
> > +		}
> >  		get_sig_err_item(sig_err_cqe, &sig->err_item);
> >  		sig->sig_err_exists = true;
> >  		sig->sigerr_count++;
> > --
> > 2.9.4
> >

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-29  1:21 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-19 10:08 [PATCH] RDMA/mlx5: Check SIG_ERR CQE mkey lookup before dereferencing lirongqing
2026-09-19 10:17 ` sashiko-bot
2026-09-28 18:38 ` Leon Romanovsky
2026-09-29  1:21   ` 答复: [????] " Li,Rongqing

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox