Linux Documentation
 help / color / mirror / Atom feed
From: Jakub Kicinski <kuba@kernel.org>
To: illusion.wang@nebula-matrix.com
Cc: Jakub Kicinski <kuba@kernel.org>,
	dimon.zhao@nebula-matrix.com, alvin.wang@nebula-matrix.com,
	sam.chen@nebula-matrix.com, netdev@vger.kernel.org,
	andrew+netdev@lunn.ch, corbet@lwn.net, horms@kernel.org,
	linux-doc@vger.kernel.org, pabeni@redhat.com,
	vadim.fedorenko@linux.dev, lukas.bulwahn@redhat.com,
	edumazet@google.com, enelsonmoore@gmail.com,
	skhan@linuxfoundation.org, hkallweit1@gmail.com,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v22 net-next 10/12] net/nebula-matrix: dispatch: add mutual exclusion lock for shared hardware resource ops
Date: Thu, 30 Jul 2026 18:30:35 -0700	[thread overview]
Message-ID: <20260731013035.745239-1-kuba@kernel.org> (raw)
In-Reply-To: <20260723040110.91410-11-illusion.wang@nebula-matrix.com>

This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
---
--- Patch [10]: [PATCH v22 net-next 10/12] net/nebula-matrix: dispatch: add mutual exclusion lock for shared hardware resource ops ---
commit 629ed395b64ead7623f563c4a19ea0794f996cf1
Author: illusion wang <illusion.wang@nebula-matrix.com>

net/nebula-matrix: dispatch: add mutual exclusion lock for shared hardware resource ops
    
Add ops_mutex_lock to serialize concurrent hardware-modifying dispatch
resource operations, preventing race conditions between PF local calls
and remote mailbox message handlers that manipulate MSI-X mapping and
mailbox IRQ state.

> 4. Fix cleanup ordering: unregister all channel message handlers before
>    hardware resource teardown to avoid lockdep splats and invalid mutex
>    access after destruction.

Does this cleanup ordering actually prevent the invalid mutex access?

Looking at the driver teardown path in nbl_core_remove() and
nbl_chan_remove_common(), it appears the channel workqueue is never cancelled
via teardown_queue(). 

If a remote PF sends mailbox messages while the local PF driver is being
unloaded, won't the uncancelled work item still process it?

> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> index c82c336f50e5e..175250f952506 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dispatch.c

[ ... ]

> @@ -227,14 +228,15 @@ static void nbl_disp_chan_configure_msix_map_resp(void *priv, u16 src_id,
>  	if (!res_ops->configure_msix_map) {
>  		err = NBL_CHAN_RESP_UNIMPLEMENTED;
>  	} else {
> -		ret = NBL_OPS_CALL_RET(res_ops->configure_msix_map,
> -				       (p, src_id,
> -					le16_to_cpu(param.num_net_msix),
> -					le16_to_cpu(param.num_others_msix),
> -					le16_to_cpu(param.msix_mask_en)));
> +		ret = NBL_OPS_CALL_LOCK_RET(disp_mgt,
> +					    res_ops->configure_msix_map, p,
> +					    src_id,
> +					    le16_to_cpu(param.num_net_msix),
> +					    le16_to_cpu(param.num_others_msix),
> +					    le16_to_cpu(param.msix_mask_en));

[Severity: High]
If the inflight work item runs after handler unregistration, will it attempt
to acquire the newly added ops_mutex_lock here via NBL_OPS_CALL_LOCK_RET?

Since this mutex is initialized via devm_mutex_init and destroyed
automatically on device detach, could this result in a use-after-free crash
on the destroyed lock?

[ ... ]

This is a pre-existing issue, but while looking at the MSI-X resource
operations I noticed a potential problem in nbl_res_intr_configure_msix_map()
in drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_interrupt.c:

release_net_bits:
        while (i--) {
                intr_index = tmp_interrupts[i];
                ...
        }
        i = num_net_msix;
release_all_bits:
        while (i--) {

[Severity: High]
Does this error path intentionally fall through to release_all_bits?

When devm_kcalloc or find_first_zero_bit fails during net interrupt
allocation, the code jumps to release_net_bits. Without a goto free_prealloc
before the i = num_net_msix assignment, it falls through into release_all_bits.

Because tmp_interrupts is unpopulated and contains zeros, will this loop
unconditionally clear bit 0 in interrupt_others_bitmap, corrupting the
MSI-X allocator state?

  reply	other threads:[~2026-07-31  1:30 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-23  4:00 [PATCH v22 net-next 00/12] nbl driver for Nebulamatrix NICs illusion.wang
2026-07-23  4:00 ` [PATCH v22 net-next 01/12] net/nebula-matrix: add minimum nbl build framework illusion.wang
2026-07-31  1:10   ` Jakub Kicinski
2026-07-23  4:00 ` [PATCH v22 net-next 02/12] net/nebula-matrix: add core driver architecture and HW layer initialization illusion.wang
2026-07-30  9:18   ` 回复:[PATCH " Illusion Wang
2026-07-31  1:13   ` [PATCH " Jakub Kicinski
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:00 ` [PATCH v22 net-next 03/12] net/nebula-matrix: add channel wire opcode enum definitions illusion.wang
2026-07-23  4:00 ` [PATCH v22 net-next 04/12] net/nebula-matrix: add channel layer illusion.wang
2026-07-31  1:27   ` Jakub Kicinski
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:00 ` [PATCH v22 net-next 05/12] net/nebula-matrix: add common resource implementation illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:00 ` [PATCH v22 net-next 06/12] net/nebula-matrix: add intr " illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:00 ` [PATCH v22 net-next 07/12] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:01 ` [PATCH v22 net-next 08/12] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
2026-07-23  4:01 ` [PATCH v22 net-next 09/12] net/nebula-matrix: dispatch: add cross-version channel message framework illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:01 ` [PATCH v22 net-next 10/12] net/nebula-matrix: dispatch: add mutual exclusion lock for shared hardware resource ops illusion.wang
2026-07-31  1:30   ` Jakub Kicinski [this message]
2026-07-23  4:01 ` [PATCH v22 net-next 11/12] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-23  4:01 ` [PATCH v22 net-next 12/12] net/nebula-matrix: add common dev start/stop operation illusion.wang
2026-07-31  1:30   ` Jakub Kicinski
2026-07-31  1:29 ` [PATCH v22 net-next 00/12] nbl driver for Nebulamatrix NICs Jakub Kicinski

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260731013035.745239-1-kuba@kernel.org \
    --to=kuba@kernel.org \
    --cc=alvin.wang@nebula-matrix.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=corbet@lwn.net \
    --cc=dimon.zhao@nebula-matrix.com \
    --cc=edumazet@google.com \
    --cc=enelsonmoore@gmail.com \
    --cc=hkallweit1@gmail.com \
    --cc=horms@kernel.org \
    --cc=illusion.wang@nebula-matrix.com \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lukas.bulwahn@redhat.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sam.chen@nebula-matrix.com \
    --cc=skhan@linuxfoundation.org \
    --cc=vadim.fedorenko@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox