Rust for Linux List
 help / color / mirror / Atom feed
From: "Gary Guo" <gary@garyguo.net>
To: "Alexandre Courbot" <acourbot@nvidia.com>, "Gary Guo" <gary@garyguo.net>
Cc: "Zhi Wang" <zhiw@nvidia.com>, <rust-for-linux@vger.kernel.org>,
	<linux-pci@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<dakr@kernel.org>, <aliceryhl@google.com>, <bhelgaas@google.com>,
	<kwilczynski@kernel.org>, <ojeda@kernel.org>, <boqun@kernel.org>,
	<bjorn3_gh@protonmail.com>, <lossin@kernel.org>,
	<a.hindborg@kernel.org>, <tmgross@umich.edu>,
	<markus.probst@posteo.de>, <cjia@nvidia.com>, <smitra@nvidia.com>,
	<ankita@nvidia.com>, <aniketa@nvidia.com>, <kwankhede@nvidia.com>,
	<targupta@nvidia.com>, <kjaju@nvidia.com>, <alkumar@nvidia.com>,
	<joelagnelf@nvidia.com>, <jhubbard@nvidia.com>,
	<zhiwang@kernel.org>
Subject: Re: [PATCH v6 1/1] rust: pci: add extended capability and SR-IOV support
Date: Fri, 31 Jul 2026 13:38:32 +0100	[thread overview]
Message-ID: <DKCRGMZJXO6X.1SWGREC78DH0W@garyguo.net> (raw)
In-Reply-To: <DKCO1B3HUGD1.1IP1J7ZZ5MHCH@nvidia.com>

On Fri Jul 31, 2026 at 10:57 AM BST, Alexandre Courbot wrote:
> On Fri Jul 31, 2026 at 3:45 AM JST, Gary Guo wrote:
>> On Thu Jul 30, 2026 at 7:29 PM BST, Zhi Wang wrote:
>>> Rust PCI drivers have no typed interface for locating and accessing PCIe
>>> extended capabilities.
>>>
>>> The SR-IOV extended capability describes VF topology and VF BARs. Expose
>>> this information through the Rust PCI abstraction so drivers can use the
>>> existing typed configuration-space accessors instead of raw bindings.
>>>
>>> Define ExtCapability to associate a capability ID with a register layout,
>>> and add ConfigSpace::find_ext_capability() to locate and project that
>>> layout. Bound the view at the next capability or the end of extended
>>> configuration space. Add ExtSriovRegs and a decoded VF BAR helper that
>>> returns the address, width, and next configuration-space slot, using
>>> PCI_SRIOV_NUM_BARS for the number of BAR slots. Since PCI_EXT_CAP_NEXT()
>>> is a function-like macro, expose it through a Rust helper.
>>>
>>> Link: https://lore.kernel.org/rust-for-linux/20260730180349.771719-1-zhiw@nvidia.com/
>>> Cc: Alexandre Courbot <acourbot@nvidia.com>
>>> Cc: Gary Guo <gary@garyguo.net>
>>> Signed-off-by: Zhi Wang <zhiw@nvidia.com>
>>> ---
>>>  rust/helpers/pci.c     |   5 +
>>>  rust/kernel/pci.rs     |   8 ++
>>>  rust/kernel/pci/cap.rs | 236 +++++++++++++++++++++++++++++++++++++++++
>>>  3 files changed, 249 insertions(+)
>>>  create mode 100644 rust/kernel/pci/cap.rs
>>>
>>> diff --git a/rust/kernel/pci/cap.rs b/rust/kernel/pci/cap.rs
>>> new file mode 100644
>>> index 000000000000..08c044bedb70
>>> --- /dev/null
>>> +++ b/rust/kernel/pci/cap.rs
>>> @@ -0,0 +1,236 @@
>>> +// SPDX-License-Identifier: GPL-2.0
>>> +
>>> +//! PCI extended capability support.
>>> +
>>> +use super::{
>>> +    io::ConfigSpaceBackend,
>>> +    ConfigSpace,
>>> +    Extended, //
>>> +};
>>> +use crate::{
>>> +    bindings,
>>> +    io::{
>>> +        Io,
>>> +        IoBackend,
>>> +        Region, //
>>> +    },
>>> +    prelude::*,
>>> +};
>>> +
>>> +/// Number of VF BAR register slots in an SR-IOV capability.
>>> +// CAST: `PCI_SRIOV_NUM_BARS` is 6, which fits in `usize`.
>>> +const NUM_VF_BARS: usize = bindings::PCI_SRIOV_NUM_BARS as usize;
>>> +
>>> +/// Attribute bits encoded in the low DWORD of a memory BAR.
>>> +const VF_MEMORY_BAR_ATTRIBUTE_BITS: u32 = bindings::PCI_BASE_ADDRESS_SPACE
>>> +    | bindings::PCI_BASE_ADDRESS_MEM_TYPE_MASK
>>> +    | bindings::PCI_BASE_ADDRESS_MEM_PREFETCH;
>>> +
>>> +/// PCI extended capability IDs.
>>> +#[repr(u16)]
>>> +#[derive(Debug, Clone, Copy, PartialEq, Eq)]
>>> +pub enum ExtCapId {
>>> +    /// Single Root I/O Virtualization.
>>> +    // CAST: `PCI_EXT_CAP_ID_SRIOV` is `0x10`, which fits in `u16`.
>>> +    Sriov = bindings::PCI_EXT_CAP_ID_SRIOV as u16,
>>> +}
>>> +
>>> +impl ExtCapId {
>>> +    fn as_raw(self) -> u16 {
>>> +        self as u16
>>> +    }
>>> +}
>>> +
>>> +/// A typed PCI extended capability register layout.
>>> +///
>>> +/// Implementors describe the register layout of one extended capability. The layout must start at
>>> +/// the extended capability header, and [`Self::ID`] must identify that layout.
>>> +pub trait ExtCapability: FromBytes + IntoBytes {
>>> +    /// PCI extended capability ID for this register layout.
>>> +    const ID: ExtCapId;
>>> +}
>>> +
>>> +impl<'a> ConfigSpace<'a, Extended> {
>>> +    /// Finds and projects an extended capability into its typed register layout.
>>> +    ///
>>> +    /// # Examples
>>> +    ///
>>> +    /// ```no_run
>>> +    /// use kernel::pci;
>>> +    ///
>>> +    /// fn probe_sriov(
>>> +    ///     pdev: &pci::Device<kernel::device::Bound>,
>>> +    /// ) -> Result<(), kernel::error::Error> {
>>> +    ///     let sriov = pdev
>>> +    ///         .config_space_extended()?
>>> +    ///         .find_ext_capability::<pci::ExtSriovRegs>()?;
>>> +    ///
>>> +    ///     let total_vfs = kernel::io_read!(sriov, .total_vfs);
>>> +    ///     let vf_offset = kernel::io_read!(sriov, .vf_offset);
>>> +    ///     let bar0 = sriov.read_vf_bar(0)?;
>>> +    ///     let bar1 = sriov.read_vf_bar(bar0.next_index())?;
>>> +    ///
>>> +    ///     Ok(())
>>> +    /// }
>>> +    /// ```
>>> +    pub fn find_ext_capability<C: ExtCapability>(&self) -> Result<ConfigSpace<'a, C>> {
>>> +        let offset = usize::from(
>>> +            // SAFETY: `self.pdev` is valid by the type invariant of `ConfigSpace`.
>>> +            unsafe {
>>> +                bindings::pci_find_ext_capability(self.pdev.as_raw(), i32::from(C::ID.as_raw()))
>>> +            },
>>> +        );
>>> +
>>> +        if offset == 0 {
>>> +            return Err(ENODEV);
>>> +        }
>>> +
>>> +        let size = self.calculate_ext_cap_size(offset);
>>> +
>>> +        let base = ConfigSpaceBackend::as_ptr(*self)
>>> +            .cast::<u8>()
>>> +            .wrapping_add(offset);
>>> +        let ptr = Region::<0>::ptr_try_from_raw_parts_mut(base, size)?;
>>
>> The signature should be 
>>
>>     Result<Option<...>>
>>
>> where the result is usually handled via `?` and `None` needs to be handled
>> explicitly, rather than matching on ENODEV.
>>
>>> +
>>> +        // SAFETY: `offset` was returned by `pci_find_ext_capability`, and
>>> +        // `calculate_ext_cap_size` bounds `ptr` at the next capability or the end of the extended
>>> +        // configuration space. `ptr_try_from_raw_parts_mut` verified the region layout.
>>> +        let capability = unsafe { ConfigSpaceBackend::project_view(*self, ptr) };
>>> +
>>> +        capability.try_cast::<C>()
>>> +    }
>>> +
>>> +    /// Calculates the size of the extended capability at `offset`.
>>> +    ///
>>> +    /// The capability extends to the next extended capability, or to the end of the extended
>>> +    /// configuration space if it is the last one. `offset` must be a DWORD-aligned offset within
>>> +    /// the extended configuration space returned by `pci_find_ext_capability`. If its header
>>> +    /// cannot be read, the capability is treated as the last one.
>>> +    fn calculate_ext_cap_size(&self, offset: usize) -> usize {
>>> +        let header = self.try_read32(offset).unwrap_or(0);
>>> +        // SAFETY: Pure bit manipulation, no preconditions.
>>> +        // CAST: The next-cap pointer is a 12-bit field (max 0xFFC), always fits in `usize`.
>>> +        let next = unsafe { bindings::pci_ext_cap_next(header) } as usize;
>>> +
>>> +        if next > offset {
>>> +            next - offset
>>> +        } else {
>>> +            (*self).size() - offset
>>> +        }
>>> +    }
>>> +}
>>> +
>>> +/// SR-IOV register layout per PCIe spec (64 bytes starting at cap offset).
>>> +#[repr(C)]
>>> +#[derive(FromBytes, IntoBytes)]
>>> +pub struct ExtSriovRegs {
>>> +    /// Extended capability header.
>>> +    pub header: u32,
>>> +    /// SR-IOV capabilities.
>>> +    pub cap: u32,
>>> +    /// SR-IOV control.
>>> +    pub ctrl: u16,
>>> +    /// SR-IOV status.
>>> +    pub status: u16,
>>> +    /// Initial VFs.
>>> +    pub initial_vfs: u16,
>>> +    /// Total VFs.
>>> +    pub total_vfs: u16,
>>> +    /// Number of VFs.
>>> +    pub num_vfs: u16,
>>> +    /// Function dependency link.
>>> +    pub func_dep_link: u8,
>>> +    _reserved_0: u8,
>>> +    /// First VF offset.
>>> +    pub vf_offset: u16,
>>> +    /// VF stride.
>>> +    pub vf_stride: u16,
>>> +    _reserved_1: u16,
>>> +    /// VF device ID.
>>> +    pub vf_device_id: u16,
>>> +    /// Supported page sizes.
>>> +    pub supported_page_sizes: u32,
>>> +    /// System page size.
>>> +    pub system_page_size: u32,
>>> +    /// VF BARs (BAR0–BAR5).
>>> +    pub vf_bar: [u32; NUM_VF_BARS],
>>> +    /// VF migration state array offset.
>>> +    pub migration_state: u32,
>>> +}
>>> +
>>> +impl ExtCapability for ExtSriovRegs {
>>> +    const ID: ExtCapId = ExtCapId::Sriov;
>>> +}
>>> +
>>> +/// A typed view of an SR-IOV extended capability.
>>> +pub type ExtSriovCapability<'a> = ConfigSpace<'a, ExtSriovRegs>;
>>> +
>>> +/// A decoded VF memory BAR.
>>> +#[derive(Debug, Clone, Copy, PartialEq, Eq)]
>>> +pub struct ExtSriovVfBar {
>>> +    address: u64,
>>> +    is_64bit: bool,
>>> +    next_index: usize,
>>> +}
>>> +
>>> +impl ExtSriovVfBar {
>>> +    /// Returns the BAR address without PCI attribute bits.
>>> +    #[inline]
>>> +    pub fn address(&self) -> u64 {
>>> +        self.address
>>> +    }
>>> +
>>> +    /// Returns whether the BAR is 64-bit.
>>> +    #[inline]
>>> +    pub fn is_64bit(&self) -> bool {
>>> +        self.is_64bit
>>> +    }
>>> +
>>> +    /// Returns the configuration-space slot index of the next logical BAR.
>>> +    #[inline]
>>> +    pub fn next_index(&self) -> usize {
>>> +        self.next_index
>>> +    }
>>> +}
>>> +
>>> +impl ConfigSpace<'_, ExtSriovRegs> {
>>> +    /// Reads and decodes the VF memory BAR at configuration-space slot `bar_index`.
>>> +    #[inline]
>>> +    pub fn read_vf_bar(&self, bar_index: usize) -> Result<ExtSriovVfBar> {
>>
>> Do you expect people to pass in a random index instead of 0 or next_index? If
>> not, this should be an iterator. Otherwise it'd be possible to index into high
>> part of 64-bit address.
>
> I made a similar suggestion on patch 3 [1], and I think Zhi kept the
> index for ergonomic reasons. But given that not all indices within range
> are valid, an iterator indeed sounds safer.
>
> [1] https://lore.kernel.org/all/DHRTUAF52GNI.1J98TSAG1LS6Q@nvidia.com/
>
>>
>>> +        if bar_index >= NUM_VF_BARS {
>>> +            return Err(EINVAL);
>>> +        }
>>> +
>>> +        let low = crate::io_read!(*self, .vf_bar[try: bar_index]);
>>
>> Given the bound checking above I'd use `panic: ` here.
>
> Or better, one could just remove the `if bar_index >= ...` block, and
> keep the `try:`. That way it will be used for what is was designed for,
> the bounds checking will be done against the actual size of the array
> and not a constant that also happens to be used as the array size, and a
> dedicated error code will be returned instead of the ubiquitous
> `EINVAL`.

`projection::OutOfBound` converts to `ERANGE` so it'll be different. Also, when
this is converted to iterators, for the first case you'd probably want to just
return `None` and for the second case log an error that the device is having out
of spec capability.

Best,
Gary


  reply	other threads:[~2026-07-31 12:38 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-30 18:29 [PATCH v6 0/1] Rust PCI capability infrastructure and SR-IOV support Zhi Wang
2026-07-30 18:29 ` [PATCH v6 1/1] rust: pci: add extended capability " Zhi Wang
2026-07-30 18:45   ` Gary Guo
2026-07-31  9:32     ` Zhi Wang
2026-07-31  9:57     ` Alexandre Courbot
2026-07-31 12:38       ` Gary Guo [this message]
2026-07-31 12:52         ` Alexandre Courbot
2026-07-31 10:35   ` Alexandre Courbot
2026-08-04  9:11     ` Zhi Wang
2026-08-09  2:28       ` Alexandre Courbot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DKCRGMZJXO6X.1SWGREC78DH0W@garyguo.net \
    --to=gary@garyguo.net \
    --cc=a.hindborg@kernel.org \
    --cc=acourbot@nvidia.com \
    --cc=aliceryhl@google.com \
    --cc=alkumar@nvidia.com \
    --cc=aniketa@nvidia.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=bjorn3_gh@protonmail.com \
    --cc=boqun@kernel.org \
    --cc=cjia@nvidia.com \
    --cc=dakr@kernel.org \
    --cc=jhubbard@nvidia.com \
    --cc=joelagnelf@nvidia.com \
    --cc=kjaju@nvidia.com \
    --cc=kwankhede@nvidia.com \
    --cc=kwilczynski@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=lossin@kernel.org \
    --cc=markus.probst@posteo.de \
    --cc=ojeda@kernel.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=smitra@nvidia.com \
    --cc=targupta@nvidia.com \
    --cc=tmgross@umich.edu \
    --cc=zhiw@nvidia.com \
    --cc=zhiwang@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox