NVIDIA GPU driver infrastructure
 help / color / mirror / Atom feed
From: "Alexandre Courbot" <acourbot@nvidia.com>
To: "John Hubbard" <jhubbard@nvidia.com>
Cc: "Danilo Krummrich" <dakr@kernel.org>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Gary Guo" <gary@garyguo.net>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them
Date: Mon, 31 Aug 2026 14:06:59 +0900	[thread overview]
Message-ID: <DL2V9T0YINMS.F1BKR8WX5T8H@nvidia.com> (raw)
In-Reply-To: <20260829013324.499542-13-jhubbard@nvidia.com>

On Sat Aug 29, 2026 at 10:33 AM JST, John Hubbard wrote:
> The GSP posts unsolicited messages onto the same queue that carries
> command replies: logs, OS error and robust-channel records, and
> lifecycle notices.
>
> Anything that was not the reply a caller awaited was discarded, and an
> unrecognized function code aborted the in-flight command, so the GSP's
> error reports never reached the log.
>
> Route every non-reply message to a dispatcher, which logs the error
> records and leaves the in-flight command waiting for its reply. The
> dispatch runs on the existing command and wait loops, so events are
> handled during normal operation before any interrupt exists. Event
> payloads, such as XID numbers and log contents, are not decoded.

The naming looks a bit misleading to me - `dispatch_event` doesn't
really dispatch anything, it logs or ignores what is passed to it. If we
plan on implementing a real dispatch mechanism in the future I'm ok with
keeping the name, but we should document that intent in its doccomment.

>
> Assisted-by: Cursor:claude-opus-5
> Signed-off-by: John Hubbard <jhubbard@nvidia.com>
> ---
>  drivers/gpu/nova-core/gsp/cmdq.rs | 64 ++++++++++++++++++++++++-------
>  1 file changed, 51 insertions(+), 13 deletions(-)
>
> diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
> index f0f28b6ded7a..0df52df1da89 100644
> --- a/drivers/gpu/nova-core/gsp/cmdq.rs
> +++ b/drivers/gpu/nova-core/gsp/cmdq.rs
> @@ -547,11 +547,11 @@ fn notify_gsp(bar: Bar0<'_>) {
>  
>      /// Sends `command` to the GSP and waits for the reply.
>      ///
> -    /// Messages with non-matching function codes are silently consumed until the expected reply
> -    /// arrives.
> +    /// A message read while waiting that is not the reply goes to
> +    /// [`CmdqInner::dispatch_event`].

`CmdqInner` is an internal private type and an implementation detail, we
shouldn't reference it in public documentation. Let's say something like
"A message read while waiting that is logged if it is an error record,
and ignored otherwise".

>      ///
> -    /// The queue is locked for the entire send+receive cycle to ensure that no other command can
> -    /// be interleaved.
> +    /// The queue is locked for the entire send+receive cycle, so no other command can be
> +    /// interleaved.

nit: is this hunk necessary?

>      ///
>      /// # Errors
>      ///
> @@ -805,8 +805,10 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
>  
>      /// Receive a message from the GSP.
>      ///
> -    /// The expected message type is specified using the `M` generic parameter. If the pending
> -    /// message has a different function code, `ERANGE` is returned and the message is consumed.
> +    /// The expected message type is specified using the `M` generic parameter. A message whose
> +    /// function code matches is decoded and returned. Any other message, whether its function code
> +    /// is a different one or is unrecognized, goes to [`Self::dispatch_event`] and `ERANGE` is
> +    /// returned.
>      ///
>      /// The read pointer is always advanced past the message, regardless of whether it matched.
>      ///
> @@ -815,8 +817,7 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
>      /// - `ETIMEDOUT` if `timeout` has elapsed before any message becomes available.
>      /// - `EIO` if there was some inconsistency (e.g. message shorter than advertised) on the
>      ///   message queue.
> -    /// - `EINVAL` if the function code of the message was not recognized.
> -    /// - `ERANGE` if the message had a recognized but non-matching function code.
> +    /// - `ERANGE` if the message was not the awaited reply.
>      ///
>      /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
>      fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
> @@ -825,11 +826,13 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
>          Error: From<M::InitError>,
>      {
>          let message = self.wait_for_msg(timeout)?;
> -        let function = message.header.function().map_err(|_| EINVAL)?;
> +        let function = message.header.function();
> +        let seq = message.header.sequence();
> +        let matched = matches!(function, Ok(f) if f == M::FUNCTION);
>  
> -        // Extract the message. Store the result as we want to advance the read pointer even in
> -        // case of failure.
> -        let result = if function == M::FUNCTION {
> +        // Bind the result rather than returning early. The read pointer must advance past this
> +        // message on every path.

Oooh nice catch, the queue was irreversibly unusable after an unknown
message is received on the previous code.

> +        let result = if matched {
>              let (cmd, contents_1) = M::Message::from_bytes_prefix(message.contents.0).ok_or(EIO)?;
>              let mut sbuffer = SBufferIter::new_reader([contents_1, message.contents.1]);
>  
> @@ -840,7 +843,7 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
>                          dev_warn!(
>                              &self.dev,
>                              "GSP message {:?} has unprocessed data\n",
> -                            function
> +                            M::FUNCTION
>                          );
>                      }
>                  })
> @@ -853,6 +856,41 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
>              message.header.length().div_ceil(GSP_PAGE_SIZE),
>          )?);
>  
> +        if !matched {
> +            self.dispatch_event(function, seq);
> +        }

We already have an `else` branch for the `if matched` statement above,
you can move this block there and avoid testing again.

It's also arguably (slightly) better since the message is now logged
before the fallible `u32::try_from` is called.

> +
>          result
>      }
> +
> +    /// Routes a GSP message that is not the reply a caller is waiting for.
> +    ///
> +    /// GSP-reported errors are logged at error level and unrecognized function codes at warning
> +    /// level. Every other known function code is consumed without a log line, because the RPC
> +    /// receive trace in [`Self::wait_for_msg`] already records its arrival.
> +    fn dispatch_event(&self, function: Result<MsgFunction, u32>, seq: u32) {

That's where we could describe the intent to turn this into an actual
dispatcher if we plan to do so in the future. Otherwise, something like
`classify_event` is probably a more fit name.

  reply	other threads:[~2026-08-31  5:07 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-29  1:22 [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:22 ` [PATCH v2 01/15] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 02/15] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-08-31  1:10   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 03/15] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-08-29  1:22 ` [PATCH v2 04/15] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-08-31 14:24   ` Alexandre Courbot
2026-09-01 13:16   ` Alexandre Courbot
2026-08-29  1:22 ` [PATCH v2 05/15] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-01  1:15   ` Alexandre Courbot
2026-08-29  1:25 ` [PATCH v2 00/15] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-08-29  1:35   ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 06/15] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-01  7:03   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 07/15] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-01 12:52   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 08/15] gpu: nova-core: dispatch GSP events instead of discarding them John Hubbard
2026-08-31  5:06   ` Alexandre Courbot [this message]
2026-08-29  1:33 ` [PATCH v2 09/15] gpu: nova-core: match GSP RPC replies by sequence, not just function John Hubbard
2026-08-31  1:09   ` Alexandre Courbot
2026-08-31  4:33     ` John Hubbard
2026-08-31 22:18       ` John Hubbard
2026-08-31 22:46         ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 10/15] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-08-31  5:35   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 11/15] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-08-31  6:04   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 12/15] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-01 14:54   ` Alexandre Courbot
2026-09-01 15:08     ` Danilo Krummrich
2026-09-02 14:33   ` Alexandre Courbot
2026-09-03  3:06     ` John Hubbard
2026-08-29  1:33 ` [PATCH v2 13/15] gpu: nova-core: retrigger the GSP falcon and clear every latched cause John Hubbard
2026-09-02 15:00   ` Alexandre Courbot
2026-08-29  1:33 ` [PATCH v2 14/15] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-08-29  1:33 ` [PATCH v2 15/15] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
2026-09-02 15:07   ` Alexandre Courbot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DL2V9T0YINMS.F1BKR8WX5T8H@nvidia.com \
    --to=acourbot@nvidia.com \
    --cc=a.hindborg@kernel.org \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=gary@garyguo.net \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox