AMD-GFX Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>
To: "Christian König" <christian.koenig@amd.com>
Cc: Geramy Loveless <gloveless@jqluv.com>,
	Bjorn Helgaas <bhelgaas@google.com>,
	alexander.deucher@amd.com, amd-gfx@lists.freedesktop.org,
	 mario.limonciello@amd.com, nra3088@gmail.com,
	 meiling.leung@embeddedllm.com, linux-pci@vger.kernel.org,
	 LKML <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH] PCI: Reserve prefetchable window headroom for Resizable BARs
Date: Mon, 31 Aug 2026 19:58:23 +0300 (EEST)	[thread overview]
Message-ID: <d6c56ff3-0b83-4d2e-7e6f-f3e65a284341@linux.intel.com> (raw)
In-Reply-To: <d4c3ea40-d92a-49c4-ab2c-925fc70c4bb9@amd.com>

[-- Attachment #1: Type: text/plain, Size: 14958 bytes --]

On Mon, 31 Aug 2026, Christian König wrote:

> On 8/28/26 23:37, Geramy Loveless wrote:
> > Firmware typically sizes prefetchable bridge windows for the boot-time
> > BAR size. Behind a fixed (non-hotplug) PCIe switch fabric, there is then
> > no room for a driver to grow a Resizable BAR afterwards: every window
> > from the leaf up to the root is sized for the small BAR, so
> > pci_resize_resource() fails with -ENOSPC. 
> > 
> > Furthermore, a small prefetchable BAR (e.g., a 2 MiB doorbell) sharing a 
> > bridge's single prefetchable window with a much larger one (e.g., a 32 GiB 
> > VRAM BAR) pushes the required window size past the large BAR's alignment. 
> > Because bridge windows round up to a power of two, this forces a massive 
> > alignment waste (e.g., 32 GiB + 2 MiB rounds up to a 64 GiB window, 
> > wasting ~32 GiB per GPU).
> > 
> > This patch solves both issues to enable ReBAR on cascaded switch fabrics:
> > 
> > 1. Reserve Headroom: 
> > Reserve prefetchable window headroom for the maximum size of each 
> > downstream Resizable BAR during the bridge sizing pass. The device BAR 
> > and hardware ReBAR are left at their boot size to prevent tearing down 
> > firmware-loaded state (e.g., AMD R9700 PSP) before the driver binds.
> 
> Yeah that was suggested before but that is clearly not something you can 
> do in common code.
> 
> The ReBAR fields often doesn't reflect the actual needed space but 
> rather the maximum the HW address logic can resolve.
> 
> So what you end up with is allocating multiple TiB for a window which 
> just needs few GiB, sometimes even completely overflowing the 64bit 
> address space made available by the root complex.

Yes. A naive approach to (only) go to the max ReBAR allows just doesn't 
work well enough to be usable in general case.

A general algorithm to pick the largests fitting sizes for ReBARs is still 
under works, a few steps away still, and more if I discover more 
challenges I need address before I can take that final step.

> > 2. Pack Oversized Windows (BAR Demotion): 
> > The PCI-to-PCI Bridge spec (r1.2, sec 3.2.5) permits a prefetchable BAR 
> > to be assigned from the non-prefetchable window. By clearing the PREFETCH 
> > flag on small control BARs during enumeration, they are placed below 4 GiB. 
> > The prefetchable window is then cleanly sized to the large BAR alone.
> 
> Interesting hack, shouldn't really matter for AMD GPUs but that is 
> something more on the heuristic side. 
> 
> A perfectly valid trick which is done by BIOS implementations but what 
> the Linux PCI subsystem still hasn't learned are back to back 
> allocations, e.g. something like this:
> 
> BAR0 32 GiB of GPU #1
> BAR1 2	MiB of GPU #1
> 
> BAR1 2	MiB of GPU #2
> BAR0 32 GiB of GPU #2

Assuming you're oversimplifying the layout above, this was added by the commit 
9036bd0efcb6 ("PCI: Align head space better").

--
 i.

> BAR0 32 GiB of GPU #3
> ...
> 
> I strongly suggest to implement that one first and see if it helps with your use case.
> 
> Regards,
> Christian.
> 
> > 
> > Why this is safe for 32-bit (< 4G) MMIO space:
> > To prevent exhausting legacy < 4G space, demotion is strictly bounded. It 
> > only triggers beside a massive prefetchable BAR (>= 64 MiB), and the demoted 
> > BAR is strictly capped at 16 MiB (PCI_BAR_PACK_CAP). Even on an 8x GPU 
> > system, this consumes at most 128 MiB of < 4G space. 
> > 
> > Fallback / Alignment logic:
> > If the 32-bit space is exhausted, small BARs remain in the prefetchable 
> > window. The alignment logic gracefully handles this by rounding the bridge 
> > window up to the next power of two (e.g., growing to 64 GiB to fit a 32 GiB 
> > BAR + BAR2) to maintain PCIe compatibility, ensuring sibling GPUs land in 
> > properly aligned slots.
> > 
> > Reserving is safe by default: __assign_resources_sorted() satisfies all
> > required resources first. Use pci=no_rebar_reserve and pci=no_bar_demote 
> > to restore previous behaviours.
> > 
> > Testing Context:
> > - ASUS K14PG-D24
> > - 2x EPYC 9354
> > - 8x Radeon AI PRO R9700 (Navi 48, 1002:7551)
> > - Broadcom PEX890xx switches
> > - Kernel base: 45c13f3f9e3b (tested with SEV-SNP passthrough=1)
> > 
> > Signed-off-by: Geramy Loveless <gloveless@jqluv.com>
> > ---
> >  .../admin-guide/kernel-parameters.txt         |   8 ++
> >  drivers/pci/pci.c                             |  10 ++
> >  drivers/pci/pci.h                             |   1 +
> >  drivers/pci/setup-bus.c                       | 118 ++++++++++++++++++
> >  4 files changed, 137 insertions(+)
> > 
> > diff --git a/Documentation/admin-guide/kernel-parameters.txt
> > b/Documentation/admin-guide/kernel-parameters.txt
> > index 37006fc3eac3..82d2b340fb98 100644
> > --- a/Documentation/admin-guide/kernel-parameters.txt
> > +++ b/Documentation/admin-guide/kernel-parameters.txt
> > @@ -5246,6 +5246,14 @@ Kernel parameters
> >  		hpbussize=nn	The minimum amount of additional bus numbers
> >  				reserved for buses below a hotplug bridge.
> >  				Default is 1.
> > +		no_rebar_reserve	Do not reserve prefetchable bridge-window
> > +				space for the maximum size of downstream Resizable
> > +				BARs. By default such space is reserved so a driver
> > +				can grow a BAR later (e.g. GPU VRAM BAR) even behind
> > +				fixed PCIe switch fabrics that firmware sized for the
> > +				boot-time BAR; the device BAR is left at its boot size
> > +				for the driver to resize and only bridge windows are
> > +				enlarged. Use this to restore the old behaviour.
> >  		realloc=	Enable/disable reallocating PCI bridge resources
> >  				if allocations done by BIOS are too small to
> >  				accommodate resources required by all child
> > diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
> > index b2879a6be5f8..c9901b15432a 100644
> > --- a/drivers/pci/pci.c
> > +++ b/drivers/pci/pci.c
> > @@ -120,6 +120,14 @@ unsigned long pci_hotplug_mmio_pref_size =
> > DEFAULT_HOTPLUG_MMIO_PREF_SIZE;
> >  #define DEFAULT_HOTPLUG_BUS_SIZE	1
> >  unsigned long pci_hotplug_bus_size = DEFAULT_HOTPLUG_BUS_SIZE;
> > 
> > +/*
> > + * Prefetchable bridge-window headroom is reserved for the maximum size of
> > + * downstream Resizable BARs by default, so a driver can grow one later
> > + * (e.g. amdgpu VRAM BAR) even behind fixed PCIe switch fabrics.
> > + * pci=no_rebar_reserve disables this.
> > + */
> > +bool pci_rebar_no_reserve;
> > +
> > 
> >  /* PCIe MPS/MRRS strategy; can be overridden by kernel command-line
> > param */
> >  enum pcie_bus_config_types pcie_bus_config = PCIE_BUS_DEFAULT;
> > @@ -6840,6 +6848,8 @@ static int __init pci_setup(char *str)
> >  					simple_strtoul(str + 10, &str, 0);
> >  				if (pci_hotplug_bus_size > 0xff)
> >  					pci_hotplug_bus_size = DEFAULT_HOTPLUG_BUS_SIZE;
> > +			} else if (!strncmp(str, "no_rebar_reserve", 16)) {
> > +				pci_rebar_no_reserve = true;
> >  			} else if (!strncmp(str, "pcie_bus_tune_off", 17)) {
> >  				pcie_bus_config = PCIE_BUS_TUNE_OFF;
> >  			} else if (!strncmp(str, "pcie_bus_safe", 13)) {
> > diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
> > index ba3c3fddddc2..6eddcb492425 100644
> > --- a/drivers/pci/pci.h
> > +++ b/drivers/pci/pci.h
> > @@ -409,6 +409,7 @@ extern unsigned long pci_hotplug_io_size;
> >  extern unsigned long pci_hotplug_mmio_size;
> >  extern unsigned long pci_hotplug_mmio_pref_size;
> >  extern unsigned long pci_hotplug_bus_size;
> > +extern bool pci_rebar_no_reserve;
> > 
> >  static inline bool pci_is_cardbus_bridge(struct pci_dev *dev)
> >  {
> > diff --git a/drivers/pci/setup-bus.c b/drivers/pci/setup-bus.c
> > index e8c94aa1d3c1..aa02b8454269 100644
> > --- a/drivers/pci/setup-bus.c
> > +++ b/drivers/pci/setup-bus.c
> > @@ -19,6 +19,7 @@
> >  #include <linux/bug.h>
> >  #include <linux/init.h>
> >  #include <linux/kernel.h>
> > +#include <linux/log2.h>
> >  #include <linux/minmax.h>
> >  #include <linux/module.h>
> >  #include <linux/pci.h>
> > @@ -1258,6 +1259,53 @@ static bool pbus_size_mem_optional(struct pci_dev
> > *dev, int resno,
> >  	return true;
> >  }
> > 
> > +/*
> > + * pci_rebar_reserve_size - prefetchable window headroom to reserve for
> > a BAR
> > + *
> > + * Reserve enough prefetchable window space to later grow @resno of @dev to
> > + * its maximum Resizable BAR size. This is done for every ReBAR-capable
> > device
> > + * by default (disable with pci=no_rebar_reserve). Only the optional
> > (add_size)
> > + * part of the enclosing window is inflated here; the device BAR
> > resource and
> > + * the hardware ReBAR are left at their boot size, so the driver still
> > performs
> > + * the actual resize into the reserved window. Reserving the window up
> > front
> > + * avoids -ENOSPC in pci_resize_resource() on fixed switch fabrics, where a
> > + * window shared with a sibling's still-assigned BAR cannot be released and
> > + * regrown at resize time.
> > + *
> > + * The reservation is strictly optional: __assign_resources_sorted()
> > satisfies
> > + * all required resources first and only then places add_size on leftover
> > + * space, so reserving by default cannot make a required BAR or window
> > fail.
> > + *
> > + * Return the extra bytes to reserve and raise *add_align to the BAR's
> > + * alignment so the reserved space can actually hold the grown BAR.
> > + */
> > +static resource_size_t pci_rebar_reserve_size(struct pci_dev *dev, int
> > resno,
> > +					      resource_size_t *add_align)
> > +{
> > +	struct resource *res = pci_resource_n(dev, resno);
> > +	resource_size_t max_size, cur_size;
> > +	int max;
> > +
> > +	if (pci_rebar_no_reserve || resno >= PCI_STD_NUM_BARS)
> > +		return 0;
> > +
> > +	if ((res->flags & (IORESOURCE_MEM_64 | IORESOURCE_PREFETCH)) !=
> > +	    (IORESOURCE_MEM_64 | IORESOURCE_PREFETCH))
> > +		return 0;
> > +
> > +	max = pci_rebar_get_max_size(dev, resno);
> > +	if (max < 0)
> > +		return 0;
> > +
> > +	max_size = pci_rebar_size_to_bytes(max);
> > +	cur_size = resource_size(res);
> > +	if (max_size <= cur_size)
> > +		return 0;
> > +
> > +	*add_align = max(*add_align, max_size);
> > +	return max_size - cur_size;
> > +}
> > +
> >  /**
> >   * pbus_size_mem() - Size the memory window of a given bus
> >   *
> > @@ -1285,6 +1333,7 @@ static void pbus_size_mem(struct pci_bus *bus,
> > struct resource *b_res,
> >  	int order, max_order;
> >  	resource_size_t children_add_size = 0;
> >  	resource_size_t add_align = 0;
> > +	resource_size_t rebar_align = 0;
> > 
> >  	if (!b_res)
> >  		return;
> > @@ -1343,15 +1392,29 @@ static void pbus_size_mem(struct pci_bus *bus,
> > struct resource *b_res,
> >  				aligns[order] += align;
> >  			if (order > max_order)
> >  				max_order = order;
> > +
> > +			size += pci_rebar_reserve_size(dev, i, &rebar_align);
> >  		}
> >  	}
> > 
> >  	win_align = pci_min_window_alignment(bus, b_res->flags);
> >  	min_align = calculate_head_align(aligns, max_order);
> >  	min_align = max(min_align, win_align);
> > +	min_align = max(min_align, rebar_align);
> >  	size0 = calculate_memsize(size, realloc_head ? 0 : add_size,
> >  				  0, win_align);
> > 
> > +	/*
> > +	 * A window that reserves ReBAR headroom is larger than its own
> > +	 * alignment (e.g. 32 GiB BAR + doorbell => 33 GiB > 32 GiB). The
> > +	 * parent only reserves alignment padding for a child window when
> > +	 * its size <= its alignment, so round such a window's alignment up
> > +	 * to a power of two >= its size; sibling 32 GiB BARs behind one
> > +	 * switch then each land in a properly aligned slot.
> > +	 */
> > +	if (rebar_align && size0)
> > +		min_align = max(min_align, roundup_pow_of_two(size0));
> > +
> >  	if (size0) {
> >  		resource_set_range(b_res, min_align, size0);
> >  		b_res->flags &= ~IORESOURCE_DISABLED;
> > @@ -2179,6 +2242,52 @@ static void pci_prepare_next_assign_round(struct
> > list_head *fail_head,
> >   * Second and later try will clear small leaf bridge res.
> >   * Will stop till to the max depth if can not find good one.
> >   */
> > +/*
> > + * Release the 64-bit prefetchable BARs of resizable-BAR display
> > devices so the
> > + * windows enclosing them empty out and can be re-sized. The BAR *size*
> > is left
> > + * untouched (only released and re-placed), so the hardware ReBAR is never
> > + * programmed here - the driver still owns the actual resize.
> > + */
> > +static int pci_release_rebar_bars_cb(struct pci_dev *dev, void *data)
> > +{
> > +	struct resource *r;
> > +	unsigned int i;
> > +
> > +	if ((dev->class >> 16) != PCI_BASE_CLASS_DISPLAY)
> > +		return 0;
> > +	if (pci_rebar_get_max_size(dev, 0) <= pci_rebar_get_current_size(dev, 0))
> > +		return 0;
> > +	pci_dev_for_each_resource(dev, r, i) {
> > +		if (i >= PCI_BRIDGE_RESOURCES)
> > +			break;
> > +		if (r->parent && (r->flags & IORESOURCE_MEM_64) &&
> > +		    (r->flags & IORESOURCE_PREFETCH))
> > +			pci_release_resource(dev, i);
> > +	}
> > +	return 0;
> > +}
> > +
> > +/*
> > + * Release the now-empty prefetchable bridge windows bottom-up so the
> > sizing
> > + * pass re-sizes them (with pci_rebar_pref_reserve() reservation) to
> > fit the
> > + * BARs a driver will later grow.
> > + */
> > +static void pci_release_rebar_windows(struct pci_bus *bus)
> > +{
> > +	struct pci_dev *dev;
> > +	struct resource *w;
> > +
> > +	list_for_each_entry(dev, &bus->devices, bus_list)
> > +		if (dev->subordinate)
> > +			pci_release_rebar_windows(dev->subordinate);
> > +
> > +	if (!bus->self)
> > +		return;
> > +	w = &bus->self->resource[PCI_BRIDGE_PREF_MEM_WINDOW];
> > +	if (w->parent && (w->flags & IORESOURCE_MEM_64) && !w->child)
> > +		pci_release_resource(bus->self, PCI_BRIDGE_PREF_MEM_WINDOW);
> > +}
> > +
> >  void pci_assign_unassigned_root_bus_resources(struct pci_bus *bus)
> >  {
> >  	LIST_HEAD(realloc_head);
> > @@ -2190,6 +2299,15 @@ void
> > pci_assign_unassigned_root_bus_resources(struct pci_bus *bus)
> >  	int pci_try_num = 1;
> >  	enum enable_type enable_local;
> > 
> > +	/*
> > +	 * Release resizable device BARs and their prefetchable windows so the
> > +	 * sizing pass re-sizes those windows large enough
> > (pci_rebar_pref_reserve)
> > +	 * for the BARs a driver will later grow. Only released and re-placed
> > - the
> > +	 * BAR size is left untouched, so the hardware ReBAR is never programmed.
> > +	 */
> > +	pci_walk_bus(bus, pci_release_rebar_bars_cb, NULL);
> > +	pci_release_rebar_windows(bus);
> > +
> >  	/* Don't realloc if asked to do so */
> >  	enable_local = pci_realloc_detect(bus, pci_realloc_enable);
> >  	if (pci_realloc_enabled(enable_local)) {
> > 
> > base-commit: 45c13f3f9e3bb15fd89ff2864c6f627a3b4b4229
> 

  reply	other threads:[~2026-09-01  7:59 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28 21:37 [PATCH] PCI: Reserve prefetchable window headroom for Resizable BARs Geramy Loveless
2026-08-31  7:49 ` Christian König
2026-08-31 16:58   ` Ilpo Järvinen [this message]
2026-08-31 17:49     ` Geramy Loveless
2026-08-31 18:07       ` Christian König
2026-08-31 18:19         ` Mario Limonciello
2026-08-31 18:32       ` Ilpo Järvinen
2026-08-31 17:52     ` Christian König
2026-08-31 17:55       ` Mario Limonciello
2026-08-31 16:31 ` Ilpo Järvinen
2026-08-31 17:40   ` Geramy Loveless
2026-08-31 18:47     ` Ilpo Järvinen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d6c56ff3-0b83-4d2e-7e6f-f3e65a284341@linux.intel.com \
    --to=ilpo.jarvinen@linux.intel.com \
    --cc=alexander.deucher@amd.com \
    --cc=amd-gfx@lists.freedesktop.org \
    --cc=bhelgaas@google.com \
    --cc=christian.koenig@amd.com \
    --cc=gloveless@jqluv.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mario.limonciello@amd.com \
    --cc=meiling.leung@embeddedllm.com \
    --cc=nra3088@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox