NVIDIA GPU driver infrastructure
 help / color / mirror / Atom feed
From: "Gary Guo" <gary@garyguo.net>
To: "Alexandre Courbot" <acourbot@nvidia.com>, "Gary Guo" <gary@garyguo.net>
Cc: "John Hubbard" <jhubbard@nvidia.com>,
	"Danilo Krummrich" <dakr@kernel.org>,
	"Timur Tabi" <ttabi@nvidia.com>,
	"Alistair Popple" <apopple@nvidia.com>,
	"Eliot Courtney" <ecourtney@nvidia.com>,
	"Zhi Wang" <zhiw@nvidia.com>, "David Airlie" <airlied@gmail.com>,
	"Simona Vetter" <simona@ffwll.ch>,
	"Bjorn Helgaas" <bhelgaas@google.com>,
	"Miguel Ojeda" <ojeda@kernel.org>,
	"Alex Gaynor" <alex.gaynor@gmail.com>,
	"Boqun Feng" <boqun.feng@gmail.com>,
	"Björn Roy Baron" <bjorn3_gh@protonmail.com>,
	"Benno Lossin" <lossin@kernel.org>,
	"Andreas Hindborg" <a.hindborg@kernel.org>,
	"Alice Ryhl" <aliceryhl@google.com>,
	"Trevor Gross" <tmgross@umich.edu>,
	nova-gpu@lists.linux.dev, LKML <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v3 10/14] gpu: nova-core: bound a GSP wait by a single deadline
Date: Fri, 04 Sep 2026 14:41:57 +0100	[thread overview]
Message-ID: <DL6KQ9S2P8AL.3BU96K936R8PZ@garyguo.net> (raw)
In-Reply-To: <DL6KJ0KDZZSA.2N71SIYKCHIGI@nvidia.com>

On Fri Sep 4, 2026 at 2:32 PM BST, Alexandre Courbot wrote:
> On Fri Sep 4, 2026 at 8:26 PM JST, Gary Guo wrote:
>> On Fri Sep 4, 2026 at 12:13 PM BST, Alexandre Courbot wrote:
>>> On Thu Sep 3, 2026 at 12:15 PM JST, John Hubbard wrote:
>>> <...>
>>>> @@ -611,15 +621,39 @@ pub(crate) fn send_command_no_wait<M>(&self, bar: Bar0<'_>, command: M) -> Resul
>>>>          self.inner.lock().send_command(bar, command)
>>>>      }
>>>>  
>>>> -    /// Receive a message from the GSP.
>>>> +    /// Waits for an unsolicited GSP event of type `M`, logging any other event that arrives
>>>> +    /// first.
>>>> +    ///
>>>> +    /// The queue is locked for the whole wait, for up to [`Self::RECEIVE_TIMEOUT`], so a
>>>> +    /// concurrent command cannot consume the awaited event.
>>>>      ///
>>>> -    /// See [`CmdqInner::receive_msg`] for details.
>>>> -    pub(crate) fn receive_msg<M: MessageFromGsp>(&self, timeout: Delta) -> Result<M>
>>>> +    /// # Errors
>>>> +    ///
>>>> +    /// - `ETIMEDOUT` if the event does not arrive within [`Self::RECEIVE_TIMEOUT`] of the call,
>>>> +    ///   however many other events arrive while waiting.
>>>> +    /// - `EIO` if the queue is poisoned or a message fails framing or checksum validation (see
>>>> +    ///   [`CmdqInner::wait_for_msg`]).
>>>
>>> Let's not mention private methods in public documentation.
>>>
>>>> +    ///
>>>> +    /// Error codes returned by [`MessageFromGsp::read`] are propagated as-is.
>>>> +    pub(crate) fn await_msg<M: MessageFromGsp>(&self) -> Result<M>
>>>>      where
>>>>          // This allows all error types, including `Infallible`, to be used for `M::InitError`.
>>>>          Error: From<M::InitError>,
>>>>      {
>>>> -        self.inner.lock().receive_msg(timeout)
>>>> +        let mut inner = self.inner.lock();
>>>> +
>>>> +        let deadline = Instant::<Monotonic>::now() + Self::RECEIVE_TIMEOUT;
>>>> +        loop {
>>>> +            let remaining = deadline - Instant::<Monotonic>::now();
>>>> +            if remaining.is_negative() {
>>>> +                break Err(ETIMEDOUT);
>>>> +            }
>>>> +            match inner.receive_msg::<M>(remaining) {
>>>> +                Ok(msg) => break Ok(msg),
>>>> +                Err(ERANGE) => continue,
>>>> +                Err(e) => break Err(e),
>>>> +            }
>>>> +        }
>>>
>>> This block and the one from `send_command` are strictly identical - we
>>> should factor them out.
>>>
>>> The right place for this seems to be a new method in `CmdqInner`:
>>>
>>>     fn await_msg<M: MessageFromGsp>(&mut self) -> Result<M> ...
>>>
>>> Then this `await_msg` simply becomes:
>>>
>>>     self.inner.lock().await_msg()
>>>
>>> While `send_command` is simplified to:
>>>
>>>     let mut inner = self.inner.lock();
>>>     inner.send_command(bar, command)?;
>>>     inner.await_msg()
>>
>> Unless I misunderstand the GSP code, the unmatched message is not discarded, but
>> rather the caller returns from inner code, and drops the lock so other waiters
>> of GSP message can have a chance to take the inner lock and receive the message
>> so then get the non-matched message out of the way. So your suggestion would
>> cause `await_msg` to never complete in such cases?
>>
>> If my understanding of this is correct, then this code should just be moved to
>> the outer `receive_msg`, because all callers of it have the same loop and I
>> think it's a wanted behaviour anyway.
>
> I don't really understand what you mean here. There is no concept of
> other waiters at the moment, and `receive_msg` unconditionally advances
> the read pointer. In effect, the queue is working in a synchronous
> manner (which is the design of the queue itself, not a Nova limitation)
> so there can be only one expected reply after a message has been
> successfully sent.
>
> I think once we move to the newer firmware we will want to add more
> sophisticated message dispatchers, but for now this simple
> implementation does what we need.

I did misread the `receive_msg` code. I read `Err(ERANGE)` as `Err(ERANGE)?` so
I thought the pointer was no incremented in such case.

So everything makes sense to me now.

Best,
Gary

>
> As for my comment, please check with the code - it's really about
> factoring out a block of code without any runtime side-effect.
>
>>
>> BTW, ERANGE is a very bad error code to mean "the message had a recognized but
>> non-matching function code".
>
> Maybe we can change this to `ENOMSG`. This will need to be its own patch
> though.



  reply	other threads:[~2026-09-04 13:42 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03  3:14 [PATCH v3 00/14] nova-core: GPU interrupt support and GSP event delivery John Hubbard
2026-09-03  3:15 ` [PATCH v3 01/14] rust: pci: declare IrqType and IrqTypes with impl_flags John Hubbard
2026-09-03  3:15 ` [PATCH v3 02/14] rust: sync: completion: add wait_for_completion_timeout() John Hubbard
2026-09-03  3:15 ` [PATCH v3 03/14] gpu: nova-core: add the GIN vector and subtree newtypes John Hubbard
2026-09-03  3:15 ` [PATCH v3 04/14] gpu: nova-core: add the GIN CPU interrupt tree and MSI EOI registers John Hubbard
2026-09-03  3:15 ` [PATCH v3 05/14] gpu: nova-core: add the per-architecture GIN CPU interrupt HAL John Hubbard
2026-09-03  3:15 ` [PATCH v3 06/14] gpu: nova-core: add the GIN interrupt tree and allocate its vectors John Hubbard
2026-09-03  3:15 ` [PATCH v3 07/14] gpu: nova-core: add an interrupt delivery self-test John Hubbard
2026-09-03  3:29   ` sashiko-bot
2026-09-03  3:57     ` John Hubbard
2026-09-03  3:15 ` [PATCH v3 08/14] gpu: nova-core: log GSP events instead of discarding them John Hubbard
2026-09-03  3:15 ` [PATCH v3 09/14] gpu: nova-core: recover the GSP receive path from corrupt framing John Hubbard
2026-09-04 10:53   ` Alexandre Courbot
2026-09-04 11:17     ` Gary Guo
2026-09-04 13:45       ` Alexandre Courbot
2026-09-03  3:15 ` [PATCH v3 10/14] gpu: nova-core: bound a GSP wait by a single deadline John Hubbard
2026-09-04 11:13   ` Alexandre Courbot
2026-09-04 11:26     ` Gary Guo
2026-09-04 13:32       ` Alexandre Courbot
2026-09-04 13:41         ` Gary Guo [this message]
2026-09-03  3:15 ` [PATCH v3 11/14] gpu: nova-core: add the falcon interrupt status and routing registers John Hubbard
2026-09-03  3:15 ` [PATCH v3 12/14] gpu: nova-core: drive GSP events with the SWGEN0 interrupt John Hubbard
2026-09-03  3:28   ` sashiko-bot
2026-09-03  3:55     ` John Hubbard
2026-09-04  1:53       ` John Hubbard
2026-09-03  3:15 ` [PATCH v3 13/14] gpu: nova-core: add KUnit tests for the interrupt tree and HALs John Hubbard
2026-09-03  3:15 ` [PATCH v3 14/14] gpu: nova-core: document the GIN interrupt controller and GSP events John Hubbard

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DL6KQ9S2P8AL.3BU96K936R8PZ@garyguo.net \
    --to=gary@garyguo.net \
    --cc=a.hindborg@kernel.org \
    --cc=acourbot@nvidia.com \
    --cc=airlied@gmail.com \
    --cc=alex.gaynor@gmail.com \
    --cc=aliceryhl@google.com \
    --cc=apopple@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun.feng@gmail.com \
    --cc=dakr@kernel.org \
    --cc=ecourtney@nvidia.com \
    --cc=jhubbard@nvidia.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=nova-gpu@lists.linux.dev \
    --cc=ojeda@kernel.org \
    --cc=simona@ffwll.ch \
    --cc=tmgross@umich.edu \
    --cc=ttabi@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox