From: John Hubbard <jhubbard@nvidia.com>
To: Danilo Krummrich <dakr@kernel.org>,
Alexandre Courbot <acourbot@nvidia.com>
Cc: "Timur Tabi" <ttabi@nvidia.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Eliot Courtney" <ecourtney@nvidia.com>,
"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Miguel Ojeda" <ojeda@kernel.org>,
"Alex Gaynor" <alex.gaynor@gmail.com>,
"Boqun Feng" <boqun.feng@gmail.com>,
"Gary Guo" <gary@garyguo.net>,
"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
"Benno Lossin" <lossin@kernel.org>,
"Andreas Hindborg" <a.hindborg@kernel.org>,
"Alice Ryhl" <aliceryhl@google.com>,
"Trevor Gross" <tmgross@umich.edu>,
nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>,
"John Hubbard" <jhubbard@nvidia.com>
Subject: [PATCH v3 10/14] gpu: nova-core: bound a GSP wait by a single deadline
Date: Wed, 2 Sep 2026 20:15:09 -0700 [thread overview]
Message-ID: <20260903031514.1515905-11-jhubbard@nvidia.com> (raw)
In-Reply-To: <20260903031514.1515905-1-jhubbard@nvidia.com>
The GSP posts unsolicited events on the same queue it posts replies on,
so a caller waiting for one message logs whatever else arrives first and
reads again.
Each of those reads started a fresh five-second timeout, so a steady
stream of events extended the wait without bound. The wait also released
the queue lock between reads, so a command sent from another thread
could consume the awaited event and leave the waiter to time out.
Compute one absolute deadline when the wait begins and pass the time
remaining to each read, and hold the queue lock across the whole wait,
so the wait is bounded however many events arrive first and no other
caller can take the event it waits for.
GSP boot waits for two unsolicited events. Move that loop into a helper
so both take the same bound.
Assisted-by: Cursor:claude-opus-5
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
---
drivers/gpu/nova-core/gsp/cmdq.rs | 50 +++++++++++++++++++++-----
drivers/gpu/nova-core/gsp/commands.rs | 8 +----
drivers/gpu/nova-core/gsp/sequencer.rs | 8 +----
3 files changed, 44 insertions(+), 22 deletions(-)
diff --git a/drivers/gpu/nova-core/gsp/cmdq.rs b/drivers/gpu/nova-core/gsp/cmdq.rs
index ce4d6a111e68..a0f995faaee5 100644
--- a/drivers/gpu/nova-core/gsp/cmdq.rs
+++ b/drivers/gpu/nova-core/gsp/cmdq.rs
@@ -33,7 +33,11 @@
},
Mutex, //
},
- time::Delta,
+ time::{
+ Delta,
+ Instant,
+ Monotonic, //
+ },
transmute::{
AsBytes,
FromBytes, //
@@ -569,8 +573,9 @@ fn notify_gsp(bar: Bar0<'_>) {
///
/// # Errors
///
- /// - `ETIMEDOUT` if space does not become available to send the command, or if the reply is
- /// not received within the timeout.
+ /// - `ETIMEDOUT` if space does not become available to send the command, or if the reply does
+ /// not arrive within [`Self::RECEIVE_TIMEOUT`] of the send, however many events arrive
+ /// while waiting.
/// - `EIO` if the variable payload requested by the command has not been entirely
/// written to by its [`CommandToGsp::init_variable_payload`] method.
///
@@ -585,8 +590,13 @@ pub(crate) fn send_command<M>(&self, bar: Bar0<'_>, command: M) -> Result<M::Rep
let mut inner = self.inner.lock();
inner.send_command(bar, command)?;
+ let deadline = Instant::<Monotonic>::now() + Self::RECEIVE_TIMEOUT;
loop {
- match inner.receive_msg::<M::Reply>(Self::RECEIVE_TIMEOUT) {
+ let remaining = deadline - Instant::<Monotonic>::now();
+ if remaining.is_negative() {
+ break Err(ETIMEDOUT);
+ }
+ match inner.receive_msg::<M::Reply>(remaining) {
Ok(reply) => break Ok(reply),
Err(ERANGE) => continue,
Err(e) => break Err(e),
@@ -611,15 +621,39 @@ pub(crate) fn send_command_no_wait<M>(&self, bar: Bar0<'_>, command: M) -> Resul
self.inner.lock().send_command(bar, command)
}
- /// Receive a message from the GSP.
+ /// Waits for an unsolicited GSP event of type `M`, logging any other event that arrives
+ /// first.
+ ///
+ /// The queue is locked for the whole wait, for up to [`Self::RECEIVE_TIMEOUT`], so a
+ /// concurrent command cannot consume the awaited event.
///
- /// See [`CmdqInner::receive_msg`] for details.
- pub(crate) fn receive_msg<M: MessageFromGsp>(&self, timeout: Delta) -> Result<M>
+ /// # Errors
+ ///
+ /// - `ETIMEDOUT` if the event does not arrive within [`Self::RECEIVE_TIMEOUT`] of the call,
+ /// however many other events arrive while waiting.
+ /// - `EIO` if the queue is poisoned or a message fails framing or checksum validation (see
+ /// [`CmdqInner::wait_for_msg`]).
+ ///
+ /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
+ pub(crate) fn await_msg<M: MessageFromGsp>(&self) -> Result<M>
where
// This allows all error types, including `Infallible`, to be used for `M::InitError`.
Error: From<M::InitError>,
{
- self.inner.lock().receive_msg(timeout)
+ let mut inner = self.inner.lock();
+
+ let deadline = Instant::<Monotonic>::now() + Self::RECEIVE_TIMEOUT;
+ loop {
+ let remaining = deadline - Instant::<Monotonic>::now();
+ if remaining.is_negative() {
+ break Err(ETIMEDOUT);
+ }
+ match inner.receive_msg::<M>(remaining) {
+ Ok(msg) => break Ok(msg),
+ Err(ERANGE) => continue,
+ Err(e) => break Err(e),
+ }
+ }
}
}
diff --git a/drivers/gpu/nova-core/gsp/commands.rs b/drivers/gpu/nova-core/gsp/commands.rs
index ffc25fd8c47b..61fe93db9e7e 100644
--- a/drivers/gpu/nova-core/gsp/commands.rs
+++ b/drivers/gpu/nova-core/gsp/commands.rs
@@ -188,13 +188,7 @@ fn read(
/// Waits for GSP initialization to complete.
pub(crate) fn wait_gsp_init_done(cmdq: &Cmdq) -> Result {
- loop {
- match cmdq.receive_msg::<GspInitDone>(Cmdq::RECEIVE_TIMEOUT) {
- Ok(_) => break Ok(()),
- Err(ERANGE) => continue,
- Err(e) => break Err(e),
- }
- }
+ cmdq.await_msg::<GspInitDone>().map(|_| ())
}
/// The `GetGspStaticInfo` command.
diff --git a/drivers/gpu/nova-core/gsp/sequencer.rs b/drivers/gpu/nova-core/gsp/sequencer.rs
index bcad1421953a..e2f1da129d8f 100644
--- a/drivers/gpu/nova-core/gsp/sequencer.rs
+++ b/drivers/gpu/nova-core/gsp/sequencer.rs
@@ -343,13 +343,7 @@ pub(crate) fn run(
libos: &'a Coherent<[LibosMemoryRegionInitArgument]>,
bootloader_app_version: u32,
) -> Result {
- let seq_info = loop {
- match cmdq.receive_msg::<GspSequence>(Cmdq::RECEIVE_TIMEOUT) {
- Ok(seq_info) => break seq_info,
- Err(ERANGE) => continue,
- Err(e) => return Err(e),
- }
- };
+ let seq_info = cmdq.await_msg::<GspSequence>()?;
let sequencer = GspSequencer {
bar: ctx.bar,
--
2.55.0
next prev parent reply other threads:[~2026-09-03 3:15 UTC|newest]
Thread overview: 27+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 3:14 [PATCH v3 00/14] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-03 3:15 ` [PATCH v3 01/14] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-03 3:15 ` [PATCH v3 02/14] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-03 3:15 ` [PATCH v3 03/14] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-09-03 3:15 ` [PATCH v3 04/14] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-03 3:15 ` [PATCH v3 05/14] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-03 3:15 ` [PATCH v3 06/14] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-03 3:15 ` [PATCH v3 07/14] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-03 3:29 ` sashiko-bot
2026-09-03 3:57 ` John Hubbard
2026-09-03 3:15 ` [PATCH v3 08/14] gpu: nova-core: log GSP events instead of discarding them John Hubbard
2026-09-03 3:15 ` [PATCH v3 09/14] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-09-04 10:53 ` Alexandre Courbot
2026-09-04 11:17 ` Gary Guo
2026-09-04 13:45 ` Alexandre Courbot
2026-09-03 3:15 ` John Hubbard [this message]
2026-09-04 11:13 ` [PATCH v3 10/14] gpu: nova-core: bound a GSP wait by a single deadline Alexandre Courbot
2026-09-04 11:26 ` Gary Guo
2026-09-04 13:32 ` Alexandre Courbot
2026-09-04 13:41 ` Gary Guo
2026-09-03 3:15 ` [PATCH v3 11/14] gpu: nova-core: add the falcon interrupt status and routing registers John Hubbard
2026-09-03 3:15 ` [PATCH v3 12/14] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-03 3:28 ` sashiko-bot
2026-09-03 3:55 ` John Hubbard
2026-09-04 1:53 ` John Hubbard
2026-09-03 3:15 ` [PATCH v3 13/14] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-09-03 3:15 ` [PATCH v3 14/14] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260903031514.1515905-11-jhubbard@nvidia.com \
--to=jhubbard@nvidia.com \
--cc=a.hindborg@kernel.org \
--cc=acourbot@nvidia.com \
--cc=airlied@gmail.com \
--cc=alex.gaynor@gmail.com \
--cc=aliceryhl@google.com \
--cc=apopple@nvidia.com \
--cc=bhelgaas@google.com \
--cc=bjorn3_gh@protonmail.com \
--cc=boqun.feng@gmail.com \
--cc=dakr@kernel.org \
--cc=ecourtney@nvidia.com \
--cc=gary@garyguo.net \
--cc=linux-kernel@vger.kernel.org \
--cc=lossin@kernel.org \
--cc=nova-gpu@lists.linux.dev \
--cc=ojeda@kernel.org \
--cc=simona@ffwll.ch \
--cc=tmgross@umich.edu \
--cc=ttabi@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox