Linux PCI subsystem development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Claudiu Beznea" <claudiu.beznea@kernel.org>
Cc: linux-pci@vger.kernel.org, wsa+renesas@sang-engineering.com
Subject: Re: [PATCH v4 8/8] PCI: rzg3s-host: Add bridge::reset_root_port()
Date: Tue, 22 Sep 2026 15:05:04 +0000	[thread overview]
Message-ID: <20260922150505.686E01F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260922144629.586997-9-claudiu.beznea@kernel.org>

Thank you for your contribution! Sashiko AI review found 7 potential issue(s) to consider:
- [High] pci: lockless RMW of RZG3S_PCI_PINTRCVIE clobbers IRQ masks
- [High] pci: lockless traversal of bus->devices causes list corruption
- [High] pci: missing synchronization of event_irq triggers abort
- [High] pci: Root Port removal breaks AER recovery logic
- [High] pci: latched GIC interrupts crash controller during reset
- [High] pci: data race on host->started imbalances clocks
- [Medium] pci: synchronize_irq(0) stall when CONFIG_PCI_MSI is disabled
--

commit 58da40640a7be86e3db083244a44a0cf4c9eb5f2
Author: Claudiu Beznea <claudiu.beznea.uj@bp.renesas.com>

    PCI: rzg3s-host: Add bridge::reset_root_port()
    
    Add bridge::reset_root_port(). This allows, on some devices, performing
    reads and writes to an NVMe endpoint after a link down/link up cycle
    without failures.

> diff --git a/drivers/pci/controller/pcie-rzg3s-host.c b/drivers/pci/controller/pcie-rzg3s-host.c
> index 1c0353598b599..9cba898a9536d 100644
> --- a/drivers/pci/controller/pcie-rzg3s-host.c
> +++ b/drivers/pci/controller/pcie-rzg3s-host.c
[ ... ]
> @@ -1110,6 +1112,11 @@ static void rzg3s_pcie_link_event(struct rzg3s_pcie_host *host)
>  
>  		dev_info(host->dev, "PCIe link down, removing devices\n");
>  
> +		for_each_pci_bridge(dev, bridge->bus) {
> +			if (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT)
> +				pci_host_handle_link_down(dev);
> +		}
> +
>  		pci_lock_rescan_remove();

[Severity: High]
Does traversing bridge->bus->devices here without holding pci_bus_sem or
pci_lock_rescan_remove() risk a use-after-free?

If concurrent hotplug, AER recovery, or sysfs unbind operations modify the
list, could the traversal follow corrupted pointers?

>  		list_for_each_entry_safe_reverse(dev, tmp, &bus->devices,
>  						 bus_list)
>  			pci_stop_and_remove_bus_device(dev);

[Severity: High]
Does this code remove the Root Port device right after initiating its recovery?

When a link down event occurs, the loop above iterates over devices on the root
bus (which are the Root Ports themselves) and invokes
pci_host_handle_link_down(dev) to schedule DPC/AER recovery.

Immediately afterward, this list_for_each_entry_safe_reverse loop executes
pci_stop_and_remove_bus_device(dev) on the exact same Root Port device. Could
this destroy the Root Port and unbind its drivers just as recovery is taking
place, breaking the recovery logic and future enumeration?

[ ... ]
> @@ -1855,6 +1862,9 @@ static int rzg3s_pcie_host_stop(struct rzg3s_pcie_host *host)
>  	struct rzg3s_sysc *sysc = host->sysc;
>  	int ret;
>  
> +	if (!host->started)
> +		return 0;
> +
>  	clk_disable_unprepare(port->refclk);

[Severity: High]
Is host->started adequately protected against concurrent access?

Because host->started is checked and modified in rzg3s_pcie_host_stop() and
rzg3s_pcie_host_reset_root_port() without any locking, concurrent invocations
of sysfs reset_root_port or PM transitions could both see host->started == true.

Could both threads subsequently disable the reference clocks and assert resets,
leading to imbalanced clock use counts and broken hardware state?

[ ... ]
> +static int rzg3s_pcie_host_reset_root_port(struct pci_host_bridge *bridge,
> +					   struct pci_dev *pdev)
> +{
> +	struct rzg3s_pcie_host *host = pci_host_bridge_priv(bridge);
> +	u32 irqs;
> +	int ret;
> +
> +	/* Mask link up/down interrupts. */
> +	writel(0, host->axi + RZG3S_PCI_PEIE0);
> +
> +	/* Mask INTx and MSI interrupts. */
> +	irqs = readl_relaxed(host->axi + RZG3S_PCI_PINTRCVIE);
> +	writel(0, host->axi + RZG3S_PCI_PINTRCVIE);

[Severity: High]
Should this mask register be read and written while holding host->hw_lock?

Concurrent calls to rzg3s_pcie_intx_irq_mask() running on another CPU will
modify this register while holding hw_lock. Could reading and clearing it here
locklessly, and restoring it later, clobber dynamic mask/unmask operations
that occurred during the reset window?

> +
> +	/*
> +	 * Make sure the next operations are not disturbed by any pending
> +	 * IRQs.
> +	 */
> +	synchronize_irq(host->msi.irq);

[Severity: Medium]
Will this stall if CONFIG_PCI_MSI is disabled?

host->msi.irq is populated only if CONFIG_PCI_MSI is enabled; otherwise, it
remains 0. Synchronizing against IRQ 0 targets a core architectural interrupt
or an invalid descriptor, which causes invalid synchronization behavior and
potential stalls.

> +	for (unsigned int i = 0; i < PCI_NUM_INTX; i++)
> +		synchronize_irq(host->intx_irqs[i]);
> +
> +	ret = rzg3s_pcie_host_stop(host);

[Severity: High]
Are there missing synchronization checks for the event_irq here?

Because host->event_irq is not synchronized, if the link state changes, the
rzg3s_pcie_event_irq_thread() can execute concurrently and attempt to read PCIe
configuration or AXI registers. Accessing the controller while clocks are
disabled by rzg3s_pcie_host_stop() immediately triggers a synchronous external
abort.

[Severity: High]
Is synchronize_irq() sufficient to prevent latched GIC interrupts from
crashing the controller during reset?

Because MSIs are edge-triggered at the GIC, an interrupt that fires immediately
before the IP-level mask can remain latched in the GIC. synchronize_irq() only
waits for currently executing handlers, so the pending GIC interrupt remains
queued.

When the CPU enables local interrupts, the GIC dispatches the handler, which
attempts to read host->axi + RZG3S_PCI_PINTRCVIS while the controller clocks
are disabled by the reset, triggering a synchronous external abort. Should
disable_irq() be used instead?

[ ... ]
> +unmask_irqs:
> +	/* Unmask INTx and MSI interrupts. */
> +	writel_relaxed(irqs, host->axi + RZG3S_PCI_PINTRCVIE);

[Severity: High]
Could this lockless restore overwrite concurrent updates made by the IRQ
subsystem?

As noted above, restoring the previously read mask state locklessly clobbers
any dynamic mask/unmask operations that occurred during the reset window, which
can cause interrupt storms or lost interrupts.

> +
> +	/* Unmask link up/down interrupts. */
> +	writel(RZG3S_PCI_PEIE0_DL_UPDOWN, host->axi + RZG3S_PCI_PEIE0);
> +
> +	host->started = true;
> +
> +	return ret;
> +}

[Severity: High]
As mentioned earlier, does writing host->started without locking allow
concurrent updates to overwrite and desync the expected controller state?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260922144629.586997-1-claudiu.beznea@kernel.org?part=8

  reply	other threads:[~2026-09-22 15:05 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 14:46 [PATCH v4 0/8] PCI: rzg3s-host: Add PCIe hotplug support Claudiu Beznea
2026-09-22 14:46 ` [PATCH v4 1/8] PCI: rzg3s-host: Follow hardware manual clock/reset initialization order Claudiu Beznea
2026-09-22 14:58   ` sashiko-bot
2026-09-24 12:19   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 2/8] PCI: rzg3s-host: Fix runtime PM handling in the NOIRQ suspend/resume phase Claudiu Beznea
2026-09-22 14:53   ` sashiko-bot
2026-09-24 12:21   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 3/8] PCI: rzg3s-host: Drop nop instructions Claudiu Beznea
2026-09-22 14:54   ` sashiko-bot
2026-09-24 12:21   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 4/8] PCI: rzg3s-host: Move host configuration code together Claudiu Beznea
2026-09-22 14:56   ` sashiko-bot
2026-09-24 12:22   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 5/8] PCI: rzg3s-host: Move suspend/resume code into dedicated functions Claudiu Beznea
2026-09-22 14:53   ` sashiko-bot
2026-09-24 12:23   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 6/8] PCI: rzg3s-host: Move IRQ domain setup code Claudiu Beznea
2026-09-22 14:58   ` sashiko-bot
2026-09-24 12:25   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 7/8] PCI: rzg3s-host: Re-enumerate the bus on PCIe link-state changes Claudiu Beznea
2026-09-22 15:05   ` sashiko-bot
2026-09-24 12:26   ` Lad, Prabhakar
2026-09-22 14:46 ` [PATCH v4 8/8] PCI: rzg3s-host: Add bridge::reset_root_port() Claudiu Beznea
2026-09-22 15:05   ` sashiko-bot [this message]
2026-09-24 12:30   ` Lad, Prabhakar

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922150505.686E01F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=claudiu.beznea@kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=wsa+renesas@sang-engineering.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox