All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Alexandre Courbot" <acourbot@nvidia.com>
To: "John Hubbard" <jhubbard@nvidia.com>
Cc: "Danilo Krummrich" <dakr@kernel.org>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Gary Guo" <gary@garyguo.net>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
	"Will Pierce" <wpierce@nvidia.com>
Subject: Re: [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt
Date: Tue, 01 Sep 2026 23:54:22 +0900	[thread overview]
Message-ID: <DL42E2ZJ3I44.1KOMPGIKIJY5S@nvidia.com> (raw)
In-Reply-To: <20260829013324.499542-17-jhubbard@nvidia.com>

Hi John,

I am still looking at this one in depth, but wanted to post what I've
found so far before going to sleep.

On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
> The GSP posts events, logs and error records to the GSP-to-CPU queue and
> raises the falcon SWGEN0 output. GSP boot polls for its own
> notifications, which leaves the latch set and pending bits in the tree.
>
> nova-core drained the queue only while polling for a command reply, so
> an event sat unread until the next command was sent.
>
> Service the queue from a threaded handler on the GSP notification
> vector. The top half runs in hard interrupt context and touches only
> registers: it clears the GIN leaf, takes the falcon's SWGEN0 latch and
> rearms PCI delivery. Draining the queue takes the command-queue mutex,
> which can sleep, so the top half wakes the IRQ thread to do it.
>
> Quiesce the tree, clear the latch and rearm PCI delivery before
> registering the handler, so none of that boot state reaches it.
> Pre-Hopper MSI rearms through a configuration-space write that the tree
> drain does not perform, and an interrupt delivered before probe leaves
> delivery un-armed.
>
> Move the vector allocation out of the self-test and into probe, because
> the vectors are allocated once for the whole PCI device rather than per
> handler. The self-test and the GSP handler each take the vector for the
> subtree they service.

Could the vector allocation have been put at its final place since patch
7?

<...>
> diff --git a/drivers/gpu/nova-core/gpu.rs b/drivers/gpu/nova-core/gpu.rs
> index 589b4b210a22..932e39e012f0 100644
> --- a/drivers/gpu/nova-core/gpu.rs
> +++ b/drivers/gpu/nova-core/gpu.rs
> @@ -10,7 +10,8 @@
>      num::Bounded,
>      pci,
>      prelude::*,
> -    sizes::SizeConstants, //
> +    sizes::SizeConstants,
> +    sync::Arc, //
>  };
>  
>  use crate::{
> @@ -25,10 +26,12 @@
>      fsp::Fsp,
>      gsp::{
>          self,
> +        cmdq::Cmdq,
>          commands::GetGspStaticInfoReply,
>          Gsp,
>          GspBootContext, //
>      },
> +    irq::SubtreeVectors,
>      vgpu::VgpuManager, //
>  };
>  
> @@ -323,12 +326,27 @@ fn drop(self: Pin<&mut Self>) {
>  }
>  
>  impl<'gpu> Gpu<'gpu> {
> +    /// Returns the chipset this GPU was identified as.
> +    pub(crate) fn chipset(&self) -> Chipset {
> +        self.spec.chipset
> +    }
> +
> +    /// Returns a shared handle to the GSP command queue.
> +    pub(crate) fn cmdq(&self) -> Arc<Cmdq> {
> +        self.gsp_resources.gsp.cmdq()
> +    }

I would expect that the interrupt rework series by Danilo would allow us
to borrow a reference to the `Cmdq` for the interrupt handler. Is there
a hard reason we cannot do it and need an `Arc`? `GspInterrupt` already
borrows `Bar0`, so this should not be different.

If the `Arc` is really needed (which I doubt, but still going through
the code), the `Arc` refactor should be in its own patch to make this
one a bit lighter and easier to review.

<...>
> diff --git a/drivers/gpu/nova-core/irq/doorbell_test.rs b/drivers/gpu/nova-core/irq/doorbell_test.rs
> index 3fd8b26e135e..c9712fa1bd18 100644
> --- a/drivers/gpu/nova-core/irq/doorbell_test.rs
> +++ b/drivers/gpu/nova-core/irq/doorbell_test.rs

I would not expect `doorbell_test.rs` to be affected by this patch -
please see if this can be squashed into patch 7.

<...>
> diff --git a/drivers/gpu/nova-core/irq/gsp.rs b/drivers/gpu/nova-core/irq/gsp.rs
> new file mode 100644
> index 000000000000..6366380eef98
> --- /dev/null
> +++ b/drivers/gpu/nova-core/irq/gsp.rs
> @@ -0,0 +1,215 @@
> +// SPDX-License-Identifier: GPL-2.0
> +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
> +
> +//! GSP event (SWGEN0) interrupt handling.
> +//!
> +//! The GSP firmware raises SWGEN0 when it has posted messages in the GSP-to-CPU queue. That
> +//! signal reaches the CPU as a PCI interrupt through the GIN tree. This module provides the
> +//! threaded IRQ handler for it. The top half services the GIN leaf and the falcon SWGEN0 latch,
> +//! and the IRQ thread drains the message queue.
> +//!
> +//! See `Documentation/gpu/nova/core/interrupts.rst`.
> +
> +use kernel::{
> +    device, irq, pci,
> +    prelude::*,
> +    sync::{
> +        aref::ARef,
> +        Arc, //
> +    },
> +};
> +
> +use super::{
> +    interrupt_tree::{
> +        GinVector,
> +        LeafEnableGuard,
> +        Subtree,
> +        Tree, //
> +    },
> +    SubtreeVectors, //
> +};
> +use crate::{
> +    driver::Bar0,
> +    falcon::gsp::Gsp as GspFalcon,
> +    gpu::Chipset,
> +    gsp::cmdq::Cmdq, //
> +};
> +
> +/// Fixed GSP notification vector.
> +///
> +/// The resource manager pins the GSP SWGEN0 notification to this vector on every supported chip,

nit: "GSP-RM pins the ..." for alignment with the rest of the docs.

> +/// so nova-core uses the constant directly instead of discovering it at runtime. The leaf and bit
> +/// serviced by the handler are derived from it.
> +const GSP_INTR_0_VECTOR: GinVector = GinVector::new::<155>();
> +
> +/// Subtree carrying the GSP notification vector, and the only subtree nova-core services.
> +///
> +/// Probe allocates PCI vectors for this subtree, and the GSP handler names it as the subtree it
> +/// serves, both when it takes its vector and when it rearms.
> +pub(crate) const GSP_SUBTREE: Subtree = GSP_INTR_0_VECTOR.subtree();
> +
> +/// Clears the interrupt state that GSP boot left behind.
> +///
> +/// Disables every vector in every implemented leaf, clears the falcon's SWGEN0 latch, clears the
> +/// tree's pending bits, and rearms PCI interrupt delivery. On return no vector is enabled, so the
> +/// tree delivers nothing.
> +pub(crate) fn quiesce(bar: Bar0<'_>, chipset: Chipset, irq_type: pci::IrqType) {
> +    let tree = Tree::new(bar, chipset, irq_type, GSP_SUBTREE.into());
> +    tree.disable_all_leaves();
> +    // GSP boot consumes its notifications by polling the queue, which leaves SWGEN0 latched.
> +    // Clear it before the tree drain below, so the drain clears the tree state the clear sets.
> +    // Messages already posted raise no interrupt of their own, and the caller's queue drain
> +    // covers them.
> +    GspFalcon::clear_swgen0_intr(bar);
> +    tree.drain();
> +    // The `TOP_EN` cycle in `drain` is the rearm for the two enable-cycle methods, but pre-Hopper
> +    // MSI rearms through a configuration-space write instead. An interrupt delivered before probe
> +    // leaves delivery un-armed on that path, with no handler to have rearmed it.
> +    tree.rearm_pci_irq(GSP_SUBTREE);
> +}
> +
> +/// Threaded IRQ handler for the GSP SWGEN0 event.
> +///
> +/// The top half clears the GIN leaf and reads the falcon SWGEN0 latch. The IRQ thread drains the
> +/// GSP-to-CPU message queue, which takes the command-queue lock.
> +#[pin_data]
> +pub(crate) struct GspInterrupt<'a> {
> +    /// Borrowed BAR0, for falcon register access from interrupt context.
> +    bar: Bar0<'a>,
> +    /// The GSP command queue, drained by the IRQ thread.
> +    cmdq: Arc<Cmdq>,

As mentioned above, I strongly suspect this can be a `&'a Cmdq`. Which
would be great as it would make the whole `Arc` refactoring unnecessary.

> +    /// The GIN interrupt tree for this chipset.
> +    tree: Tree<'a>,
> +    /// Device, for logging from interrupt context without taking the command-queue lock.
> +    dev: ARef<device::Device>,

This can be a `&'a device::Device` (tested locally).

Will post some more tomorrow but this is what I have so far. :)

  reply	other threads:[~2026-09-01 14:54 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-29  1:22 [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:22 ` [PATCH v2 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 03/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-08-29  1:22 ` [PATCH v2 04/15] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-08-31 14:24   ` Alexandre Courbot
2026-09-01 13:16   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-01  1:15   ` Alexandre Courbot
2026-08-29  1:25 ` [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:35   ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-01  7:03   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 07/15] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-01 12:52   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them John Hubbard
2026-08-31  5:06   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 09/15] gpu: nova-core: match GSP RPC replies by sequence, not just function John Hubbard
2026-08-31  1:09   ` Alexandre Courbot
2026-08-31  4:33     ` John Hubbard
2026-08-31 22:18       ` John Hubbard
2026-08-31 22:46         ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-08-31  5:35   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-08-31  6:04   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-01 14:54   ` Alexandre Courbot [this message]
2026-09-01 15:08     ` Danilo Krummrich
2026-09-02 14:33   ` Alexandre Courbot
2026-09-03  3:06     ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 13/15] gpu: nova-core: retrigger the GSP falcon and clear every latched cause John Hubbard
2026-09-02 15:00   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 14/15] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-08-29  1:33 ` [PATCH v2 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
2026-09-02 15:07   ` Alexandre Courbot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DL42E2ZJ3I44.1KOMPGIKIJY5S@nvidia.com \
    --to=acourbot@nvidia.com \
    --cc=a.hindborg@kernel.org \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=wpierce@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.