From: John Hubbard <jhubbard@nvidia.com>
To: Danilo Krummrich <dakr@kernel.org>,
Alexandre Courbot <acourbot@nvidia.com>
Cc: "Timur Tabi" <ttabi@nvidia.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Eliot Courtney" <ecourtney@nvidia.com>,
"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Miguel Ojeda" <ojeda@kernel.org>,
"Alex Gaynor" <alex.gaynor@gmail.com>,
"Boqun Feng" <boqun.feng@gmail.com>,
"Gary Guo" <gary@garyguo.net>,
"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
"Benno Lossin" <lossin@kernel.org>,
"Andreas Hindborg" <a.hindborg@kernel.org>,
"Alice Ryhl" <aliceryhl@google.com>,
"Trevor Gross" <tmgross@umich.edu>,
nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
"John Hubbard" <jhubbard@nvidia.com>
Subject: [PATCH v4 09/17] gpu: nova-core: log GSP events instead of discarding them
Date: Fri, 11 Sep 2026 21:43:52 -0700 [thread overview]
Message-ID: <20260912044400.677097-10-jhubbard@nvidia.com> (raw)
In-Reply-To: <20260912044400.677097-1-jhubbard@nvidia.com>
The GSP posts unsolicited messages on the same queue that carries
command replies: log records, OS error and robust-channel records, and
lifecycle notices.
nova-core discarded every message that was not the reply a caller was
waiting for, and an unrecognized function code aborted the in-flight
command. The GSP's error reports never reached the kernel log.
Log every non-reply message according to its function code, on the
receive path that already reads it, and leave the in-flight command
waiting for its reply. Event payloads, such as XID numbers and log
contents, are not decoded.
Assisted-by: LLM
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
---
drivers/gpu/nova-core/gsp/cmdq.rs | 52 ++++++++++++++++++++++++-------
1 file changed, 41 insertions(+), 11 deletions(-)
diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
index 9f99e6bbb4fa..340760e384d0 100644
--- a/drivers/gpu/nova-core/gsp/cmdq.rs
+++ b/drivers/gpu/nova-core/gsp/cmdq.rs
@@ -557,8 +557,7 @@ fn notify_gsp(bar: Bar0<'_>) {
/// Sends `command` to the GSP and waits for the reply.
///
- /// Messages with non-matching function codes are silently consumed until the expected reply
- /// arrives.
+ /// Events that arrive before the reply are logged and consumed.
///
/// The queue is locked for the entire send+receive cycle to ensure that no other command can
/// be interleaved.
@@ -815,8 +814,8 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
/// Receive a message from the GSP.
///
- /// The expected message type is specified using the `M` generic parameter. If the pending
- /// message has a different function code, `ERANGE` is returned and the message is consumed.
+ /// A message whose function code is `M::FUNCTION` is decoded and returned. Any other message
+ /// is logged as an event.
///
/// The read pointer is always advanced past the message, regardless of whether it matched.
///
@@ -825,8 +824,7 @@ fn wait_for_msg(&self, timeout: Delta) -> Result<GspMessage<'_>> {
/// - `ETIMEDOUT` if `timeout` has elapsed before any message becomes available.
/// - `EIO` if there was some inconsistency (e.g. message shorter than advertised) on the
/// message queue.
- /// - `EINVAL` if the function code of the message was not recognized.
- /// - `ERANGE` if the message had a recognized but non-matching function code.
+ /// - `ERANGE` if the message was not the awaited reply.
///
/// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
@@ -835,11 +833,11 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
Error: From<M::InitError>,
{
let message = self.wait_for_msg(timeout)?;
- let function = message.header.function().map_err(|_| EINVAL)?;
+ let function = message.header.function();
+ let seq = message.header.sequence();
- // Extract the message. Store the result as we want to advance the read pointer even in
- // case of failure.
- let result = if function == M::FUNCTION {
+ // An early return here would leave the read pointer on this message.
+ let result = if matches!(function, Ok(f) if f == M::FUNCTION) {
let (cmd, contents_1) = M::Message::from_bytes_prefix(message.contents.0).ok_or(EIO)?;
let mut sbuffer = SBufferIter::new_reader([contents_1, message.contents.1]);
@@ -850,11 +848,13 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
dev_warn!(
&self.dev,
"GSP message {:?} has unprocessed data\n",
- function
+ M::FUNCTION
);
}
})
} else {
+ self.log_event(function, seq);
+
Err(ERANGE)
};
@@ -865,4 +865,34 @@ fn receive_msg<M: MessageFromGsp>(&mut self, timeout: Delta) -> Result<M>
result
}
+
+ /// Logs an event, meaning a message that no caller was waiting for.
+ ///
+ /// An OS error or robust-channel record is logged at error level and an unknown function code
+ /// at warning level. Every other event is recorded only by the receive trace in
+ /// [`Self::wait_for_msg`].
+ fn log_event(&self, function: Result<MsgFunction, u32>, seq: u32) {
+ match function {
+ Ok(MsgFunction::OsErrorLog) => {
+ dev_err!(&self.dev, "GSP reported an OS error (seq {})\n", seq);
+ }
+ Ok(MsgFunction::RcTriggered) => {
+ dev_err!(
+ &self.dev,
+ "GSP triggered robust-channel recovery (seq {})\n",
+ seq
+ );
+ }
+ // Nothing to do for the remaining known function codes.
+ Ok(_) => {}
+ Err(raw) => {
+ dev_warn!(
+ &self.dev,
+ "unknown GSP message function {:#x} (seq {})\n",
+ raw,
+ seq
+ );
+ }
+ }
+ }
}
--
2.55.0
next prev parent reply other threads:[~2026-09-12 4:44 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 4:43 [PATCH v4 00/17] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-12 4:43 ` [PATCH v4 01/17] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-12 4:43 ` [PATCH v4 02/17] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-12 4:43 ` [PATCH v4 03/17] gpu: nova-core: add the GIN vector, leaf and subtree types John Hubbard
2026-09-12 4:43 ` [PATCH v4 04/17] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-12 4:43 ` [PATCH v4 05/17] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-12 4:43 ` [PATCH v4 06/17] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-12 4:43 ` [PATCH v4 07/17] gpu: nova-core: wait for GFW boot in probe, not in the Gpu constructor John Hubbard
2026-09-12 4:43 ` [PATCH v4 08/17] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-12 4:43 ` John Hubbard [this message]
2026-09-12 4:43 ` [PATCH v4 10/17] gpu: nova-core: stop re-parsing a bad GSP message John Hubbard
2026-09-12 4:43 ` [PATCH v4 11/17] gpu: nova-core: return ENOMSG for an unmatched " John Hubbard
2026-09-12 4:43 ` [PATCH v4 12/17] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-09-12 4:43 ` [PATCH v4 13/17] gpu: nova-core: add a GSP message queue drain John Hubbard
2026-09-12 4:43 ` [PATCH v4 14/17] gpu: nova-core: add the falcon interrupt registers and their HAL John Hubbard
2026-09-12 4:43 ` [PATCH v4 15/17] gpu: nova-core: service GSP events from the SWGEN0 interrupt John Hubbard
2026-09-12 4:43 ` [PATCH v4 16/17] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-09-12 4:44 ` [PATCH v4 17/17] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260912044400.677097-10-jhubbard@nvidia.com \
--to=jhubbard@nvidia.com \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=aliceryhl@google.com \
--cc=apopple@nvidia.com \
--cc=bhelgaas@google.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=dakr@kernel.org \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=tmgross@umich.edu \
--cc=ttabi@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.