Linux PCI subsystem development
 help / color / mirror / Atom feed
From: Claudiu Beznea <claudiu.beznea@kernel.org>
To: sashiko-reviews@lists.linux.dev
Cc: wsa+renesas@sang-engineering.com, linux-pci@vger.kernel.org
Subject: Re: [PATCH v5 9/9] PCI: rzg3s-host: Add bridge::reset_root_port()
Date: Mon, 5 Oct 2026 17:55:12 +0300	[thread overview]
Message-ID: <595ca63d-5768-48ca-bc2e-aebb1fafd0d6@kernel.org> (raw)
In-Reply-To: <20261002112518.9DD401F000FF@smtp.kernel.org>



On 10/2/26 14:25, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
> - [High] pci: rzg3s: INTx/MSI mask state loss due to stale register snapshot
> - [High] pci: rzg3s: fatal unpowered MMIO access in PM notifier after reset failure
> - [High] pci: rzg3s: lockless iteration over bus->devices causes list corruption
> - [Medium] pci: rzg3s: missing I/O flush allows IRQs to fire after synchronize_irq
> --
> 
> commit c5ffdc83b56df877b7b137c1229cfe9487c3fb34
> Author: Claudiu Beznea <claudiu.beznea.uj@bp.renesas.com>
> Subject: PCI: rzg3s-host: Add bridge::reset_root_port()
> 
> This commit adds a reset_root_port() callback for the rzg3s-host bridge.
> It introduces the started boolean to track the controller state and prevents
> double start/stop from breaking reference counters. It also updates the link
> event and power management notifier paths to handle the interrupts correctly.
> 
>> diff --git a/drivers/pci/controller/pcie-rzg3s-host.c b/drivers/pci/controller/pcie-rzg3s-host.c
>> index 78e783928b9dc..eea215818633b 100644
>> --- a/drivers/pci/controller/pcie-rzg3s-host.c
>> +++ b/drivers/pci/controller/pcie-rzg3s-host.c
> [ ... ]
>> @@ -1115,6 +1117,11 @@ static void rzg3s_pcie_link_event(struct rzg3s_pcie_host *host)
>>   
>>   		dev_info(host->dev, "PCIe link down, removing devices\n");
>>   
>> +		for_each_pci_bridge(dev, bridge->bus) {
>> +			if (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT)
>> +				pci_host_handle_link_down(dev);
>> +		}
>> +
>>   		pci_lock_rescan_remove();
> 
> [Severity: High]
> Is it safe to iterate over bus->devices without holding the PCI
> rescan/remove lock?

Doing this will lead to deadlock on the following calls:
pci_host_handle_link_down() ->
   pcie_do_recovery() ->
     pci_host_reset_root_port() ->
       pci_lock_rescan_remove()

> 
> It looks like concurrent modifications to the list from sysfs rescan or
> hotplug events could cause list corruption or a use-after-free. The
> necessary lock is correctly taken immediately after this loop for the removal
> step, but appears to be omitted for this iteration.

I kept it as is to align with the current usage of pci_host_handle_link_down() 
in other drivers.

> 
> [ ... ]
>> @@ -1956,6 +1974,64 @@ static int rzg3s_pcie_host_start(struct rzg3s_pcie_host *host)
>>   	return ret;
>>   }
>>   
>> +static int rzg3s_pcie_host_reset_root_port(struct pci_host_bridge *bridge,
>> +					   struct pci_dev *pdev)
>> +{
>> +	struct rzg3s_pcie_host *host = pci_host_bridge_priv(bridge);
>> +	unsigned long flags;
>> +	u32 irqs;
>> +	int ret;
>> +
>> +	/* Mask link up/down interrupts. */
>> +	writel(0, host->axi + RZG3S_PCI_PEIE0);
>> +
>> +	/* Mask INTx and MSI interrupts. */
>> +	raw_spin_lock_irqsave(&host->hw_lock, flags);
>> +	irqs = readl(host->axi + RZG3S_PCI_PINTRCVIE);
>> +	writel(0, host->axi + RZG3S_PCI_PINTRCVIE);
>> +	raw_spin_unlock_irqrestore(&host->hw_lock, flags);
>> +
>> +	/*
>> +	 * Make sure the next operations are not disturbed by any pending
>> +	 * IRQs.
>> +	 */
>> +	if (IS_ENABLED(CONFIG_PCI_MSI))
>> +		synchronize_irq(host->msi.irq);
> 
> [Severity: Medium]
> Does the posted write to the interrupt mask register need to be flushed
> with a dummy read before calling synchronize_irq()?

OK

> 
> Without a read-back to ensure the mask has reached the physical hardware, the
> controller might still emit an interrupt while synchronize_irq() runs,
> potentially creating a race window where the handler crashes when the
> controller clocks are subsequently cut in rzg3s_pcie_host_stop().
> 
>> +	for (unsigned int i = 0; i < PCI_NUM_INTX; i++)
>> +		synchronize_irq(host->intx_irqs[i]);
>> +
>> +	ret = rzg3s_pcie_host_stop(host);
>> +	if (ret) {
>> +		dev_err(host->dev, "Failed to stop the host!\n");
>> +		goto unmask_irqs;
>> +	}
>> +
>> +	ret = rzg3s_pcie_host_start(host, false);
>> +	if (ret) {
>> +		dev_err(host->dev, "Failed to start the host!\n");
>> +
>> +		/*
>> +		 * Don't unmask IRQs. We are in a bad state here and we
>> +		 * can recover only through a suspend/resume cycle. Just
>> +		 * return and preserve the stop state.
>> +		 */
>> +		return ret;
>> +	}
>> +
>> +unmask_irqs:
>> +	/* Unmask INTx and MSI interrupts. */
>> +	raw_spin_lock_irqsave(&host->hw_lock, flags);
>> +	writel_relaxed(irqs, host->axi + RZG3S_PCI_PINTRCVIE);
>> +	raw_spin_unlock_irqrestore(&host->hw_lock, flags);
> 
> [Severity: High]
> Could this unconditional restore overwrite mask changes made by concurrent
> INTx mask/unmask operations?
> 
> Since hw_lock was dropped earlier to wait for IRQs and execute the
> controller stop/start sequence, any intermediate updates to the register by
> other drivers dynamically masking or unmasking their IRQs might be
> permanently lost here when the old snapshot is restored.

[1] had the mechanism to avoid this. It was discussed (not exactly on this 
particular use case) to drop it.

[1] 
https://lore.kernel.org/all/n7cafny3y6rr63tx6rkt5umxl3f4rsxpu2ojze7luudzzr4flb@egj7dru2bnhq

> 
>> +
>> +	/* Unmask link up/down interrupts. */
>> +	writel_relaxed(RZG3S_PCI_PEIE0_DL_UPDOWN, host->axi + RZG3S_PCI_PEIE0);
>> +
>> +	host->started = true;
>> +
>> +	return ret;
>> +}
> [ ... ]
>> @@ -2036,6 +2112,7 @@ static int rzg3s_pcie_pm_notifier(struct notifier_block *nb,
>>   	case PM_SUSPEND_PREPARE:
>>   		/* Disable link up/down interrupts. */
>>   		disable_irq(host->event_irq);
>> +		writel_relaxed(0, host->axi + RZG3S_PCI_PEIE0);
> 
> [Severity: High]
> Could this write cause a synchronous external abort if the controller is
> currently unpowered?
> 
> If rzg3s_pcie_host_start() fails inside rzg3s_pcie_host_reset_root_port(),
> the function intentionally returns early, leaving host->started false and
> the controller in reset with clocks disabled. A subsequent system suspend
> would unconditionally hit this AXI write, which might hang the bus since
> the target peripheral is held in reset.

Will update to take into account the host->started in both PM_SUSPEND_PREPARE 
and PM_POST_SUSPEND.

> 
>>   		break;
>>   
>>   	case PM_POST_SUSPEND:
> 


      reply	other threads:[~2026-10-05 14:55 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-02 11:09 [PATCH v5 0/9] PCI: rzg3s-host: Add PCIe hotplug support Claudiu Beznea
2026-10-02 11:09 ` [PATCH v5 1/9] PCI: rzg3s-host: Follow hardware manual clock/reset initialization order Claudiu Beznea
2026-10-02 11:23   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 2/9] PCI: rzg3s-host: Fix runtime PM handling in the NOIRQ suspend/resume phase Claudiu Beznea
2026-10-02 11:18   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 3/9] PCI: rzg3s-host: Select PCI_HOST_COMMON Claudiu Beznea
2026-10-02 11:22   ` sashiko-bot
2026-10-05 13:50     ` Claudiu Beznea
2026-10-02 11:09 ` [PATCH v5 4/9] PCI: rzg3s-host: Drop nop instructions Claudiu Beznea
2026-10-02 11:17   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 5/9] PCI: rzg3s-host: Move host configuration code together Claudiu Beznea
2026-10-02 11:20   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 6/9] PCI: rzg3s-host: Move suspend/resume code into dedicated functions Claudiu Beznea
2026-10-02 11:17   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 7/9] PCI: rzg3s-host: Move IRQ domain setup code Claudiu Beznea
2026-10-02 11:19   ` sashiko-bot
2026-10-02 11:09 ` [PATCH v5 8/9] PCI: rzg3s-host: Re-enumerate the bus on PCIe link-state changes Claudiu Beznea
2026-10-02 11:25   ` sashiko-bot
2026-10-05 13:50     ` Claudiu Beznea
2026-10-02 11:09 ` [PATCH v5 9/9] PCI: rzg3s-host: Add bridge::reset_root_port() Claudiu Beznea
2026-10-02 11:25   ` sashiko-bot
2026-10-05 14:55     ` Claudiu Beznea [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=595ca63d-5768-48ca-bc2e-aebb1fafd0d6@kernel.org \
    --to=claudiu.beznea@kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=wsa+renesas@sang-engineering.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox