From: Jakub Kicinski <kuba@kernel.org>
To: illusion.wang@nebula-matrix.com
Cc: Jakub Kicinski <kuba@kernel.org>,
dimon.zhao@nebula-matrix.com, alvin.wang@nebula-matrix.com,
sam.chen@nebula-matrix.com, netdev@vger.kernel.org,
andrew+netdev@lunn.ch, corbet@lwn.net, horms@kernel.org,
linux-doc@vger.kernel.org, pabeni@redhat.com,
vadim.fedorenko@linux.dev, lukas.bulwahn@redhat.com,
edumazet@google.com, enelsonmoore@gmail.com,
skhan@linuxfoundation.org, hkallweit1@gmail.com,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v24 net-next 11/11] net/nebula-matrix: add common dev start/stop operation
Date: Fri, 14 Aug 2026 12:48:19 -0700 [thread overview]
Message-ID: <20260814194819.3201177-1-kuba@kernel.org> (raw)
In-Reply-To: <20260807044422.2356-12-illusion.wang@nebula-matrix.com>
This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
---
net/nebula-matrix: add common dev start/stop operation
Adds nbl_dev_start()/nbl_dev_stop() to configure the hardware MSI-X map,
allocate MSI-X vectors, request/free the mailbox IRQ and enable/disable the
mailbox interrupt, and hooks nbl_core_start()/nbl_core_stop() into the PCI
probe and remove paths. Also drops the NULL checks around the chip
init_module/deinit_module ops.
A few questions below.
> net/nebula-matrix: add common dev start/stop operation
>
> Add device start/stop helper functions to manage MSI-X vector mapping,
> mailbox interrupt resource initialization and deinitialization.
> Create unbound global workqueue for mailbox task scheduling; this shared
> work infrastructure supports both current interrupt-driven receive and
> future polling receive path.
[ ... ]
> Extend channel TX waiting logic with polling fallback and shutdown detection.
> When interrupt path is disabled, send routines rely on active polling to
> wait for message ACK, preventing deadlock during device stop.
[Severity: Low]
Do these two paragraphs describe changes that are in this patch?
Diffing this commit against its parent for nbl_common/ and nbl_channel/
gives an empty diff. The workqueue is created by the pre-existing
nbl_common_create_wq():
common->wq = alloc_workqueue("%s", WQ_UNBOUND, 0, wq_name);
which is per-PCI-device rather than global, and the polling/shutdown-aware
ACK wait already exists in nbl_chan_send_msg() from the preceding
"add common/ctrl dev init/remove operation" commit.
This patch only consumes both, via queue_work() in
nbl_dev_clean_mailbox_schedule() and chan_ops->set_queue_state(). Could
the changelog be adjusted so bisection and backporting point at the commit
that actually adds them?
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> index bdb4d6a45fbc..3a8b7ee2e475 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_core/nbl_dev.c
[ ... ]
> @@ -14,6 +25,187 @@ static void nbl_dev_init_msix_cnt(struct nbl_dev_mgt *dev_mgt)
> msix_info->serv_info[NBL_MSIX_MAILBOX_TYPE].num = 1;
> }
>
> +static int nbl_dev_request_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
> +{
[ ... ]
> + snprintf(dev_common->mailbox_name, sizeof(dev_common->mailbox_name),
> + "nbl_mailbox@pci:%s", pci_name(common->pdev));
> + err = request_irq(irq_num, nbl_dev_clean_mailbox, 0,
> + dev_common->mailbox_name, dev_mgt);
> + if (err)
> + return err;
> +
> + return 0;
> +}
[ ... ]
> +static int nbl_dev_disable_mailbox_irq(struct nbl_dev_mgt *dev_mgt)
> +{
[ ... ]
> + * This helper is invoked in two paths:
> + * 1. Error unwind path of nbl_dev_start(): immediately followed by
> + * nbl_dev_free_mailbox_irq(), channel resources
> + * are fully torn down afterwards, no stale descriptors remain.
> + * 2. Normal device stop path nbl_dev_stop(): free_irq() synchronously
> + * waits for any in-flight threaded irq bottom-half handler to
> + * finish execution, all pending mailbox cleanup work completes
> + * before channel teardown in late remove stage. No stuck
> + * descriptors will linger in either scenario.
> + */
[Severity: Low]
Is there actually a threaded IRQ bottom half here for free_irq() to wait
on?
nbl_dev_request_mailbox_irq() uses request_irq(), which passes
thread_fn == NULL to request_threaded_irq(), so nbl_dev_clean_mailbox()
runs entirely in hardirq context.
The real bottom half is a workqueue item:
static irqreturn_t nbl_dev_clean_mailbox(int irq, void *data)
{
nbl_dev_clean_mailbox_schedule(dev_mgt); /* queue_work(common->wq, ...) */
return IRQ_HANDLED;
}
free_irq() has no knowledge of that work item. And nbl_dev_stop()
follows free_irq() with cancel_work_sync(&common_dev->clean_mbx_task),
which dequeues a queued-but-not-yet-started item so it never runs, rather
than letting "all pending mailbox cleanup work complete".
Could the comment be reworded to match what the code does?
> + chan_ops->set_queue_state(dev_mgt->chan_ops_tbl->priv,
> + NBL_CHAN_INTERRUPT_READY,
> + NBL_CHAN_TYPE_MAILBOX, false);
> +
> + return disp_ops->set_mailbox_irq(dev_mgt->disp_ops_tbl->priv,
> + local_vec_id, false);
> +}
[Severity: Low]
This isn't a bug today, but the wait strategy in nbl_chan_send_msg() is
picked once from an unlocked test_bit():
if (test_bit(NBL_CHAN_INTERRUPT_READY, chan_info->state)) {
... wait_event_timeout(wait_head->wait_queue,
READ_ONCE(wait_head->acked) || READ_ONCE(chan_info->shutdown),
NBL_CHAN_ACK_WAIT_TIME);
A sender that already sampled the bit as set and committed to that branch
would not be released by nbl_dev_stop(), since the stop path never sets
chan_info->shutdown (only nbl_chan_teardown_queue() does, and it runs
later) and never wakes the wait queues, so it would burn the full ack
timeout and report a spurious failure.
Enumerating the ack=1 senders, all of them are the dispatch request
wrappers called only from nbl_dev_init()/nbl_dev_start()/nbl_dev_stop(),
i.e. probe/remove for that device, which the driver core serialises, and
the mailbox RX path only sends acks with ack=0. So no second waiter can
exist for a given chan_info at the moment. Is it worth re-checking the
bit inside the wait, or setting shutdown/waking the waiters here, for
future concurrent users?
[ ... ]
> @@ -238,3 +438,62 @@ void nbl_dev_remove(struct nbl_adapter *adapter)
[ ... ]
> +/* ---------- Dev start process ---------- */
> +int nbl_dev_start(struct nbl_adapter *adapter)
> +{
> + struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
> + struct nbl_dev_common *common_dev = dev_mgt->common_dev;
> + int cleanup_ret;
> + int ret;
> +
> + ret = nbl_dev_cfg_msix_map(dev_mgt);
> + if (ret)
> + return ret;
[Severity: Medium]
Should this return jump to init_interrupt_scheme_err instead, so that
nbl_dev_destroy_msix_map() runs?
Every other step in the sequence has an unwind label, and this one can
fail after the hardware side already committed. On a non-control PF
(has_ctrl == 0, i.e. any function other than PF0), cfg_msix_map is
nbl_disp_chan_cfg_msix_map_req(), a mailbox RPC that returns whatever
nbl_chan_send_msg() returns, including -ETIMEDOUT when the ack is lost or
late, while the control PF responder already ran res_ops->cfg_msix_map()
successfully.
On the control PF that call reaches the point of no return:
/* Phase2: All new resource allocation succeeded. */
ret = nbl_res_intr_destroy_msix_map(res_mgt, func_id);
...
intr_mgt->func_intr_res[func_id].interrupts = tmp_interrupts;
after which the kcalloc'ed interrupts[] array, the dmam_alloc_coherent
MSI-X map table and the bits set in intr_net_bmap/intr_other_bmap are only
released by nbl_res_intr_destroy_msix_map(), reachable only through
destroy_msix_map.
Combined with the probe error path below clearing drvdata, nbl_dev_stop()
never runs afterwards, so destroy_msix_map is never issued. Does this
leak the control PF's interrupts[] array, the MSI-X map DMA table and the
vector bitmap bits?
> +
> + ret = nbl_dev_init_interrupt_scheme(dev_mgt);
> + if (ret)
> + goto init_interrupt_scheme_err;
[ ... ]
> +void nbl_dev_stop(struct nbl_adapter *adapter)
> +{
> + struct nbl_dev_mgt *dev_mgt = adapter->core.dev_mgt;
> + struct nbl_dev_common *common_dev = dev_mgt->common_dev;
> + int ret;
> +
> + ret = nbl_dev_disable_mailbox_irq(dev_mgt);
> + if (ret)
> + dev_err(dev_mgt->common->dev,
> + "Failed to disable mailbox IRQ: %d\n", ret);
> + nbl_dev_free_mailbox_irq(dev_mgt);
> + cancel_work_sync(&common_dev->clean_mbx_task);
> + nbl_dev_clear_interrupt_scheme(dev_mgt);
> + ret = nbl_dev_destroy_msix_map(dev_mgt);
[Severity: Low]
Does this cancel_work_sync() achieve anything, given what the next call
does?
nbl_dev_disable_mailbox_irq() has just cleared NBL_CHAN_INTERRUPT_READY,
so on a non-control PF nbl_dev_destroy_msix_map() ->
nbl_disp_chan_destroy_msix_map_req() -> send_msg(..., ack=1) takes the
polling branch of nbl_chan_send_msg(), which re-arms the same work item on
each poll iteration:
task = READ_ONCE(chan_info->clean_task);
if (common->wq && task && !READ_ONCE(chan_info->shutdown))
queue_work(common->wq, task);
clean_task is &common_dev->clean_mbx_task, registered by
nbl_dev_register_chan_task(), so clean_mbx_task is pending or running
again by the time nbl_dev_stop() returns.
It is harmless today because nbl_dev_remove_chan_queue() ->
nbl_chan_teardown_queue() later drains inflight senders, cancel_work_sync()s
clean_task and NULLs it before nbl_common_destroy_wq(). Would moving the
cancel after destroy_msix_map(), or dropping it, be clearer?
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
> index bd175e200408..c1e6180a7010 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_hw/nbl_chip.c
> @@ -9,8 +9,7 @@ void nbl_res_chip_deinit_module(struct nbl_resource_mgt *res_mgt)
> {
> struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
>
> - if (hw_ops->deinit_module)
> - hw_ops->deinit_module(res_mgt->hw_ops_tbl->priv);
> + hw_ops->deinit_module(res_mgt->hw_ops_tbl->priv);
> }
>
> int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt)
> @@ -20,7 +19,5 @@ int nbl_res_chip_init_module(struct nbl_resource_mgt *res_mgt)
> struct nbl_hw_ops *hw_ops = res_mgt->hw_ops_tbl->ops;
> struct nbl_hw_mgt *p = res_mgt->hw_ops_tbl->priv;
>
> - if (hw_ops->init_module)
> - return hw_ops->init_module(p, eth_speed, eth_num);
> - return -EOPNOTSUPP;
> + return hw_ops->init_module(p, eth_speed, eth_num);
> }
[Severity: Low]
Is this hunk related to dev start/stop? Nothing in the changelog mentions
removing the NULL guards or the -EOPNOTSUPP fallback.
The removal looks safe today, since the single hw_ops instance statically
initialises both callbacks and nbl_hw_setup_ops() rejects a table with
NULL init_module/deinit_module, but none of that is visible from the
changelog. Could this go into its own patch, or at least be mentioned?
> diff --git a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> index 8a968a06b5b6..c12a53da6a9b 100644
> --- a/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> +++ b/drivers/net/ethernet/nebula-matrix/nbl/nbl_main.c
> @@ -131,7 +141,14 @@ static int nbl_probe(struct pci_dev *pdev,
> goto adapter_init_err;
> }
> pci_set_drvdata(pdev, adapter);
> + err = nbl_core_start(adapter);
> + if (err)
> + goto core_start_err;
> +
> return 0;
> +core_start_err:
> + pci_set_drvdata(pdev, NULL);
> + nbl_core_remove(adapter);
> adapter_init_err:
> pci_clear_master(pdev);
> configure_dma_err:
This is the path referenced above: clearing drvdata here means nbl_remove()
sees a NULL adapter, so nbl_core_stop() and therefore
nbl_dev_destroy_msix_map() can never run for a probe that failed inside
nbl_dev_start().
prev parent reply other threads:[~2026-08-14 19:48 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-07 4:44 [PATCH v24 net-next 00/11] nbl driver for Nebulamatrix NICs illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 01/11] net/nebula-matrix: add minimum nbl build framework illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 02/11] net/nebula-matrix: add core driver architecture and HW layer initialization illusion.wang
2026-08-14 19:33 ` Jakub Kicinski
2026-08-07 4:44 ` [PATCH v24 net-next 03/11] net/nebula-matrix: add channel wire opcode enum definitions illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 04/11] net/nebula-matrix: add channel layer illusion.wang
2026-08-14 19:38 ` Jakub Kicinski
2026-08-07 4:44 ` [PATCH v24 net-next 05/11] net/nebula-matrix: add common resource implementation illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 06/11] net/nebula-matrix: add intr " illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 07/11] net/nebula-matrix: add chip-wide hardware init/deinit implementation illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 08/11] net/nebula-matrix: dispatch: add control-level routing core infrastructure illusion.wang
2026-08-14 19:45 ` Jakub Kicinski
2026-08-07 4:44 ` [PATCH v24 net-next 09/11] net/nebula-matrix: dispatch: add channel RPC framework & shared hw ops mutex illusion.wang
2026-08-07 4:44 ` [PATCH v24 net-next 10/11] net/nebula-matrix: add common/ctrl dev init/remove operation illusion.wang
2026-08-14 19:48 ` Jakub Kicinski
2026-08-07 4:44 ` [PATCH v24 net-next 11/11] net/nebula-matrix: add common dev start/stop operation illusion.wang
2026-08-14 19:48 ` Jakub Kicinski [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260814194819.3201177-1-kuba@kernel.org \
--to=kuba@kernel.org \
--cc=alvin.wang@nebula-matrix.com \
--cc=andrew+netdev@lunn.ch \
--cc=corbet@lwn.net \
--cc=dimon.zhao@nebula-matrix.com \
--cc=edumazet@google.com \
--cc=enelsonmoore@gmail.com \
--cc=hkallweit1@gmail.com \
--cc=horms@kernel.org \
--cc=illusion.wang@nebula-matrix.com \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lukas.bulwahn@redhat.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=sam.chen@nebula-matrix.com \
--cc=skhan@linuxfoundation.org \
--cc=vadim.fedorenko@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox