NVIDIA GPU driver infrastructure
 help / color / mirror / Atom feed
From: "Alexandre Courbot" <acourbot@nvidia.com>
To: "John Hubbard" <jhubbard@nvidia.com>
Cc: "Danilo Krummrich" <dakr@kernel.org>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Gary Guo" <gary@garyguo.net>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
	"Will Pierce" <wpierce@nvidia.com>
Subject: Re: [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt
Date: Tue, 01 Sep 2026 23:54:22 +0900	[thread overview]
Message-ID: <DL42E2ZJ3I44.1KOMPGIKIJY5S@nvidia.com> (raw)
In-Reply-To: <20260829013324.499542-17-jhubbard@nvidia.com>

Hi John,

I am still looking at this one in depth, but wanted to post what I've
found so far before going to sleep.

On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
> The GSP posts events, logs and error records to the GSP-to-CPU queue and
> raises the falcon SWGEN0 output. GSP boot polls for its own
> notifications, which leaves the latch set and pending bits in the tree.
>
> nova-core drained the queue only while polling for a command reply, so
> an event sat unread until the next command was sent.
>
> Service the queue from a threaded handler on the GSP notification
> vector. The top half runs in hard interrupt context and touches only
> registers: it clears the GIN leaf, takes the falcon's SWGEN0 latch and
> rearms PCI delivery. Draining the queue takes the command-queue mutex,
> which can sleep, so the top half wakes the IRQ thread to do it.
>
> Quiesce the tree, clear the latch and rearm PCI delivery before
> registering the handler, so none of that boot state reaches it.
> Pre-Hopper MSI rearms through a configuration-space write that the tree
> drain does not perform, and an interrupt delivered before probe leaves
> delivery un-armed.
>
> Move the vector allocation out of the self-test and into probe, because
> the vectors are allocated once for the whole PCI device rather than per
> handler. The self-test and the GSP handler each take the vector for the
> subtree they service.

Could the vector allocation have been put at its final place since patch
7?

<...>
> diff --git a/drivers/gpu/nova-core/gpu.rs b/drivers/gpu/nova-core/gpu.rs
> index 589b4b210a22..932e39e012f0 100644
> --- a/drivers/gpu/nova-core/gpu.rs
> +++ b/drivers/gpu/nova-core/gpu.rs
> @@ -10,7 +10,8 @@
>      num::Bounded,
>      pci,
>      prelude::*,
> -    sizes::SizeConstants, //
> +    sizes::SizeConstants,
> +    sync::Arc, //
>  };
>  
>  use crate::{
> @@ -25,10 +26,12 @@
>      fsp::Fsp,
>      gsp::{
>          self,
> +        cmdq::Cmdq,
>          commands::GetGspStaticInfoReply,
>          Gsp,
>          GspBootContext, //
>      },
> +    irq::SubtreeVectors,
>      vgpu::VgpuManager, //
>  };
>  
> @@ -323,12 +326,27 @@ fn drop(self: Pin<&mut Self>) {
>  }
>  
>  impl<'gpu> Gpu<'gpu> {
> +    /// Returns the chipset this GPU was identified as.
> +    pub(crate) fn chipset(&self) -> Chipset {
> +        self.spec.chipset
> +    }
> +
> +    /// Returns a shared handle to the GSP command queue.
> +    pub(crate) fn cmdq(&self) -> Arc<Cmdq> {
> +        self.gsp_resources.gsp.cmdq()
> +    }

I would expect that the interrupt rework series by Danilo would allow us
to borrow a reference to the `Cmdq` for the interrupt handler. Is there
a hard reason we cannot do it and need an `Arc`? `GspInterrupt` already
borrows `Bar0`, so this should not be different.

If the `Arc` is really needed (which I doubt, but still going through
the code), the `Arc` refactor should be in its own patch to make this
one a bit lighter and easier to review.

<...>
> diff --git a/drivers/gpu/nova-core/irq/doorbell_test.rs b/drivers/gpu/nova-core/irq/doorbell_test.rs
> index 3fd8b26e135e..c9712fa1bd18 100644
> --- a/drivers/gpu/nova-core/irq/doorbell_test.rs
> +++ b/drivers/gpu/nova-core/irq/doorbell_test.rs

I would not expect `doorbell_test.rs` to be affected by this patch -
please see if this can be squashed into patch 7.

<...>
> diff --git a/drivers/gpu/nova-core/irq/gsp.rs b/drivers/gpu/nova-core/irq/gsp.rs
> new file mode 100644
> index 000000000000..6366380eef98
> --- /dev/null
> +++ b/drivers/gpu/nova-core/irq/gsp.rs
> @@ -0,0 +1,215 @@
> +// SPDX-License-Identifier: GPL-2.0
> +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
> +
> +//! GSP event (SWGEN0) interrupt handling.
> +//!
> +//! The GSP firmware raises SWGEN0 when it has posted messages in the GSP-to-CPU queue. That
> +//! signal reaches the CPU as a PCI interrupt through the GIN tree. This module provides the
> +//! threaded IRQ handler for it. The top half services the GIN leaf and the falcon SWGEN0 latch,
> +//! and the IRQ thread drains the message queue.
> +//!
> +//! See `Documentation/gpu/nova/core/interrupts.rst`.
> +
> +use kernel::{
> +    device, irq, pci,
> +    prelude::*,
> +    sync::{
> +        aref::ARef,
> +        Arc, //
> +    },
> +};
> +
> +use super::{
> +    interrupt_tree::{
> +        GinVector,
> +        LeafEnableGuard,
> +        Subtree,
> +        Tree, //
> +    },
> +    SubtreeVectors, //
> +};
> +use crate::{
> +    driver::Bar0,
> +    falcon::gsp::Gsp as GspFalcon,
> +    gpu::Chipset,
> +    gsp::cmdq::Cmdq, //
> +};
> +
> +/// Fixed GSP notification vector.
> +///
> +/// The resource manager pins the GSP SWGEN0 notification to this vector on every supported chip,

nit: "GSP-RM pins the ..." for alignment with the rest of the docs.

> +/// so nova-core uses the constant directly instead of discovering it at runtime. The leaf and bit
> +/// serviced by the handler are derived from it.
> +const GSP_INTR_0_VECTOR: GinVector = GinVector::new::<155>();
> +
> +/// Subtree carrying the GSP notification vector, and the only subtree nova-core services.
> +///
> +/// Probe allocates PCI vectors for this subtree, and the GSP handler names it as the subtree it
> +/// serves, both when it takes its vector and when it rearms.
> +pub(crate) const GSP_SUBTREE: Subtree = GSP_INTR_0_VECTOR.subtree();
> +
> +/// Clears the interrupt state that GSP boot left behind.
> +///
> +/// Disables every vector in every implemented leaf, clears the falcon's SWGEN0 latch, clears the
> +/// tree's pending bits, and rearms PCI interrupt delivery. On return no vector is enabled, so the
> +/// tree delivers nothing.
> +pub(crate) fn quiesce(bar: Bar0<'_>, chipset: Chipset, irq_type: pci::IrqType) {
> +    let tree = Tree::new(bar, chipset, irq_type, GSP_SUBTREE.into());
> +    tree.disable_all_leaves();
> +    // GSP boot consumes its notifications by polling the queue, which leaves SWGEN0 latched.
> +    // Clear it before the tree drain below, so the drain clears the tree state the clear sets.
> +    // Messages already posted raise no interrupt of their own, and the caller's queue drain
> +    // covers them.
> +    GspFalcon::clear_swgen0_intr(bar);
> +    tree.drain();
> +    // The `TOP_EN` cycle in `drain` is the rearm for the two enable-cycle methods, but pre-Hopper
> +    // MSI rearms through a configuration-space write instead. An interrupt delivered before probe
> +    // leaves delivery un-armed on that path, with no handler to have rearmed it.
> +    tree.rearm_pci_irq(GSP_SUBTREE);
> +}
> +
> +/// Threaded IRQ handler for the GSP SWGEN0 event.
> +///
> +/// The top half clears the GIN leaf and reads the falcon SWGEN0 latch. The IRQ thread drains the
> +/// GSP-to-CPU message queue, which takes the command-queue lock.
> +#[pin_data]
> +pub(crate) struct GspInterrupt<'a> {
> +    /// Borrowed BAR0, for falcon register access from interrupt context.
> +    bar: Bar0<'a>,
> +    /// The GSP command queue, drained by the IRQ thread.
> +    cmdq: Arc<Cmdq>,

As mentioned above, I strongly suspect this can be a `&'a Cmdq`. Which
would be great as it would make the whole `Arc` refactoring unnecessary.

> +    /// The GIN interrupt tree for this chipset.
> +    tree: Tree<'a>,
> +    /// Device, for logging from interrupt context without taking the command-queue lock.
> +    dev: ARef<device::Device>,

This can be a `&'a device::Device` (tested locally).

Will post some more tomorrow but this is what I have so far. :)

  reply	other threads:[~2026-09-01 14:54 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-29  1:22 [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:22 ` [PATCH v2 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 03/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-08-29  1:22 ` [PATCH v2 04/15] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-08-31 14:24   ` Alexandre Courbot
2026-09-01 13:16   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-01  1:15   ` Alexandre Courbot
2026-08-29  1:25 ` [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:35   ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-01  7:03   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 07/15] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-01 12:52   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them John Hubbard
2026-08-31  5:06   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 09/15] gpu: nova-core: match GSP RPC replies by sequence, not just function John Hubbard
2026-08-31  1:09   ` Alexandre Courbot
2026-08-31  4:33     ` John Hubbard
2026-08-31 22:18       ` John Hubbard
2026-08-31 22:46         ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-08-31  5:35   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-08-31  6:04   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-01 14:54   ` Alexandre Courbot [this message]
2026-09-01 15:08     ` Danilo Krummrich
2026-09-02 14:33   ` Alexandre Courbot
2026-09-03  3:06     ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 13/15] gpu: nova-core: retrigger the GSP falcon and clear every latched cause John Hubbard
2026-09-02 15:00   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 14/15] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-08-29  1:33 ` [PATCH v2 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
2026-09-02 15:07   ` Alexandre Courbot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DL42E2ZJ3I44.1KOMPGIKIJY5S@nvidia.com \
    --to=acourbot@nvidia.com \
    --cc=a.hindborg@kernel.org \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=wpierce@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox