The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely
@ 2026-08-25  5:17 Sahil Chandna
  2026-08-25  6:46 ` Naman Jain
  0 siblings, 1 reply; 4+ messages in thread
From: Sahil Chandna @ 2026-08-25  5:17 UTC (permalink / raw)
  To: kys, haiyangz, wei.liu, decui, longli, lpieralisi, kwilczynski,
	mani, robh, bhelgaas, linux-hyperv, linux-pci, linux-kernel

A guest can wait indefinitely in wait_for_response() for the host to
send either a rescind message or a packet completion. If the
host does not send either, the guest can remain blocked with no
diagnostic indicating a reason.
This was observed during a guest kernel upgrade in which the
host-side application handling the PCI channel faulted, causing the
guest to never receive the completion request.
Add a periodic warning in wait_for_response() when the wait exceeds
a timeout so that such a hang is visible in the guest's kernel log
and can be correlated with host-side state.

Suggested-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
---
This was sent earlier upstream [1]
[1] https://lore.kernel.org/linux-hyperv/20260612174010.2598695-1-hamzamahfooz@linux.microsoft.com/
---
 drivers/pci/controller/pci-hyperv.c | 14 +++++++++++++-
 1 file changed, 13 insertions(+), 1 deletion(-)

diff --git a/drivers/pci/controller/pci-hyperv.c b/drivers/pci/controller/pci-hyperv.c
index cfc8fa403dad..c4fba0039164 100644
--- a/drivers/pci/controller/pci-hyperv.c
+++ b/drivers/pci/controller/pci-hyperv.c
@@ -1040,11 +1040,18 @@ static void put_pcichild(struct hv_pci_dev *hpdev)

 /*
  * There is no good way to get notified from vmbus_onoffer_rescind(),
- * so let's use polling here, since this is not a hot path.
+ * so let's use polling here, since this is not a hot path. If
+ * wait_for_response() has been polling for PCI_RESPONSE_HANG_TIMEOUT_SEC
+ * without either a rescind or completion, add a periodic warning.
  */
+#define PCI_RESPONSE_HANG_TIMEOUT_SEC 300
+
 static int wait_for_response(struct hv_device *hdev,
 			     struct completion *comp)
 {
+	unsigned long delay = secs_to_jiffies(PCI_RESPONSE_HANG_TIMEOUT_SEC);
+	u64 timeout = get_jiffies_64() + delay;
+
 	while (true) {
 		if (hdev->channel->rescind) {
 			dev_warn_once(&hdev->device, "The device is gone.\n");
@@ -1053,6 +1060,11 @@ static int wait_for_response(struct hv_device *hdev,

 		if (wait_for_completion_timeout(comp, HZ / 10))
 			break;
+
+		if (time_after64(get_jiffies_64(), timeout)) {
+			dev_warn(&hdev->device, "PCI stuck waiting for response.\n");
+			timeout = get_jiffies_64() + delay;
+		}
 	}

 	return 0;
--
2.53.0

^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely
  2026-08-25  5:17 [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely Sahil Chandna
@ 2026-08-25  6:46 ` Naman Jain
  2026-08-25 12:00   ` Sahil Chandna
  2026-08-25 17:00   ` Long Li
  0 siblings, 2 replies; 4+ messages in thread
From: Naman Jain @ 2026-08-25  6:46 UTC (permalink / raw)
  To: Sahil Chandna, kys, haiyangz, wei.liu, decui, longli, lpieralisi,
	kwilczynski, mani, robh, bhelgaas, linux-hyperv, linux-pci,
	linux-kernel



On 8/25/2026 10:47 AM, Sahil Chandna wrote:
> A guest can wait indefinitely in wait_for_response() for the host to
> send either a rescind message or a packet completion. If the
> host does not send either, the guest can remain blocked with no
> diagnostic indicating a reason.
> This was observed during a guest kernel upgrade in which the
> host-side application handling the PCI channel faulted, causing the
> guest to never receive the completion request.
> Add a periodic warning in wait_for_response() when the wait exceeds
> a timeout so that such a hang is visible in the guest's kernel log
> and can be correlated with host-side state.
> 
> Suggested-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
> Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
> ---
> This was sent earlier upstream [1]
> [1] https://lore.kernel.org/linux-hyperv/20260612174010.2598695-1-hamzamahfooz@linux.microsoft.com/
> ---
>   drivers/pci/controller/pci-hyperv.c | 14 +++++++++++++-
>   1 file changed, 13 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/pci/controller/pci-hyperv.c b/drivers/pci/controller/pci-hyperv.c
> index cfc8fa403dad..c4fba0039164 100644
> --- a/drivers/pci/controller/pci-hyperv.c
> +++ b/drivers/pci/controller/pci-hyperv.c
> @@ -1040,11 +1040,18 @@ static void put_pcichild(struct hv_pci_dev *hpdev)
> 
>   /*
>    * There is no good way to get notified from vmbus_onoffer_rescind(),
> - * so let's use polling here, since this is not a hot path.
> + * so let's use polling here, since this is not a hot path. If
> + * wait_for_response() has been polling for PCI_RESPONSE_HANG_TIMEOUT_SEC
> + * without either a rescind or completion, add a periodic warning.
>    */
> +#define PCI_RESPONSE_HANG_TIMEOUT_SEC 300
> +
>   static int wait_for_response(struct hv_device *hdev,
>   			     struct completion *comp)
>   {
> +	unsigned long delay = secs_to_jiffies(PCI_RESPONSE_HANG_TIMEOUT_SEC);
> +	u64 timeout = get_jiffies_64() + delay;
> +
>   	while (true) {
>   		if (hdev->channel->rescind) {
>   			dev_warn_once(&hdev->device, "The device is gone.\n");
> @@ -1053,6 +1060,11 @@ static int wait_for_response(struct hv_device *hdev,
> 
>   		if (wait_for_completion_timeout(comp, HZ / 10))
>   			break;
> +
> +		if (time_after64(get_jiffies_64(), timeout)) {
> +			dev_warn(&hdev->device, "PCI stuck waiting for response.\n");
> +			timeout = get_jiffies_64() + delay;
> +		}
>   	}

There can be some enhancements in above patch to address these problems:
1. Logging forever every 5 minutes in case of no completion or rescind.
2. If we now print warning once, not knowing if completion ever arrived.
3. On solving pt. 1 and 2 by adding a print for completion, one should 
avoid adding a print by default for regular timely completions.

Basically something like this:

#define PCI_RESPONSE_WARN_TIMEOUT_SEC 300

static int wait_for_response(struct hv_device *hdev,
                  struct completion *comp)
{
     unsigned long warn_at =
         jiffies + secs_to_jiffies(PCI_RESPONSE_WARN_TIMEOUT_SEC);
     bool warned = false;

     while (true) {
         if (hdev->channel->rescind) {
             dev_warn_once(&hdev->device, "The device is gone.\n");
             return -ENODEV;
         }

         if (wait_for_completion_timeout(comp, HZ / 10)) {
             if (warned || time_after_eq(jiffies, warn_at))
                 dev_warn(&hdev->device,
                      "PCI response received after prolonged wait.\n");
             return 0;
         }

         if (!warned && time_after_eq(jiffies, warn_at)) {
             dev_warn(&hdev->device,
                  "PCI still waiting for response.\n");
             warned = true;
         }
     }
}

Regards,
Naman

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely
  2026-08-25  6:46 ` Naman Jain
@ 2026-08-25 12:00   ` Sahil Chandna
  2026-08-25 17:00   ` Long Li
  1 sibling, 0 replies; 4+ messages in thread
From: Sahil Chandna @ 2026-08-25 12:00 UTC (permalink / raw)
  To: Naman Jain, kys, haiyangz, wei.liu, decui, longli, lpieralisi,
	kwilczynski, mani, robh, bhelgaas, linux-hyperv, linux-pci,
	linux-kernel

On 25-08-2026 12:16, Naman Jain wrote:
> 
> 
> On 8/25/2026 10:47 AM, Sahil Chandna wrote:
>> A guest can wait indefinitely in wait_for_response() for the host to
>> send either a rescind message or a packet completion. If the
>> host does not send either, the guest can remain blocked with no
>> diagnostic indicating a reason.
>> This was observed during a guest kernel upgrade in which the
>> host-side application handling the PCI channel faulted, causing the
>> guest to never receive the completion request.
>> Add a periodic warning in wait_for_response() when the wait exceeds
>> a timeout so that such a hang is visible in the guest's kernel log
>> and can be correlated with host-side state.
>>
>> Suggested-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
>> Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
>> ---
>> This was sent earlier upstream [1]
>> [1] https://lore.kernel.org/linux-hyperv/20260612174010.2598695-1-hamzamahfooz@linux.microsoft.com/
>> ---
>>   drivers/pci/controller/pci-hyperv.c | 14 +++++++++++++-
>>   1 file changed, 13 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/pci/controller/pci-hyperv.c b/drivers/pci/controller/pci-hyperv.c
>> index cfc8fa403dad..c4fba0039164 100644
>> --- a/drivers/pci/controller/pci-hyperv.c
>> +++ b/drivers/pci/controller/pci-hyperv.c
>> @@ -1040,11 +1040,18 @@ static void put_pcichild(struct hv_pci_dev *hpdev)
>>
>>   /*
>>    * There is no good way to get notified from vmbus_onoffer_rescind(),
>> - * so let's use polling here, since this is not a hot path.
>> + * so let's use polling here, since this is not a hot path. If
>> + * wait_for_response() has been polling for PCI_RESPONSE_HANG_TIMEOUT_SEC
>> + * without either a rescind or completion, add a periodic warning.
>>    */
>> +#define PCI_RESPONSE_HANG_TIMEOUT_SEC 300
>> +
>>   static int wait_for_response(struct hv_device *hdev,
>>   			     struct completion *comp)
>>   {
>> +	unsigned long delay = secs_to_jiffies(PCI_RESPONSE_HANG_TIMEOUT_SEC);
>> +	u64 timeout = get_jiffies_64() + delay;
>> +
>>   	while (true) {
>>   		if (hdev->channel->rescind) {
>>   			dev_warn_once(&hdev->device, "The device is gone.\n");
>> @@ -1053,6 +1060,11 @@ static int wait_for_response(struct hv_device *hdev,
>>
>>   		if (wait_for_completion_timeout(comp, HZ / 10))
>>   			break;
>> +
>> +		if (time_after64(get_jiffies_64(), timeout)) {
>> +			dev_warn(&hdev->device, "PCI stuck waiting for response.\n");
>> +			timeout = get_jiffies_64() + delay;
>> +		}
>>   	}
> 
> There can be some enhancements in above patch to address these problems:
> 1. Logging forever every 5 minutes in case of no completion or rescind.
> 2. If we now print warning once, not knowing if completion ever arrived.
> 3. On solving pt. 1 and 2 by adding a print for completion, one should 
> avoid adding a print by default for regular timely completions.
> 
Ack.
> Basically something like this:
> 
> #define PCI_RESPONSE_WARN_TIMEOUT_SEC 300
> 
> static int wait_for_response(struct hv_device *hdev,
>                   struct completion *comp)
> {
>      unsigned long warn_at =
>          jiffies + secs_to_jiffies(PCI_RESPONSE_WARN_TIMEOUT_SEC);
>      bool warned = false;
> 
>      while (true) {
>          if (hdev->channel->rescind) {
>              dev_warn_once(&hdev->device, "The device is gone.\n");
>              return -ENODEV;
>          }
> 
>          if (wait_for_completion_timeout(comp, HZ / 10)) {
>              if (warned || time_after_eq(jiffies, warn_at))
>                  dev_warn(&hdev->device,
>                       "PCI response received after prolonged wait.\n");
>              return 0;
>          }
> 
>          if (!warned && time_after_eq(jiffies, warn_at)) {
>              dev_warn(&hdev->device,
>                   "PCI still waiting for response.\n");
>              warned = true;
>          }
>      }
> }
> 
Thanks for review, I align on repeated warning every 5 minutes for
stalled guest would add to dmesg noise.I will wait for other review
comments as well and share v2 addressing this.
> Regards,
> Naman


^ permalink raw reply	[flat|nested] 4+ messages in thread

* RE: [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely
  2026-08-25  6:46 ` Naman Jain
  2026-08-25 12:00   ` Sahil Chandna
@ 2026-08-25 17:00   ` Long Li
  1 sibling, 0 replies; 4+ messages in thread
From: Long Li @ 2026-08-25 17:00 UTC (permalink / raw)
  To: Naman Jain, Sahil Chandna, KY Srinivasan, Haiyang Zhang,
	wei.liu@kernel.org, Dexuan Cui, lpieralisi@kernel.org,
	kwilczynski@kernel.org, mani@kernel.org, robh@kernel.org,
	bhelgaas@google.com, linux-hyperv@vger.kernel.org,
	linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org

> 
> 
> On 8/25/2026 10:47 AM, Sahil Chandna wrote:
> > A guest can wait indefinitely in wait_for_response() for the host to
> > send either a rescind message or a packet completion. If the host does
> > not send either, the guest can remain blocked with no diagnostic
> > indicating a reason.
> > This was observed during a guest kernel upgrade in which the host-side
> > application handling the PCI channel faulted, causing the guest to
> > never receive the completion request.
> > Add a periodic warning in wait_for_response() when the wait exceeds a
> > timeout so that such a hang is visible in the guest's kernel log and
> > can be correlated with host-side state.
> >
> > Suggested-by: Hamza Mahfooz <hamzamahfooz@linux.microsoft.com>
> > Signed-off-by: Sahil Chandna <sahilchandna@linux.microsoft.com>
> > ---
> > This was sent earlier upstream [1]
> > [1]
> > https://nam06.safelinks.protection.outlook.com/?url=https%3A%2F%2Flore
> > .kernel.org%2Flinux-hyperv%2F20260612174010.2598695-1-
> hamzamahfooz%40l
> >
> inux.microsoft.com%2F&data=05%7C02%7Clongli%40microsoft.com%7C22b60
> f06
> >
> 81624fda801f08df02748dce%7C72f988bf86f141af91ab2d7cd011db47%7C1%7
> C0%7C
> >
> 639232371774209409%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOn
> RydWUsIlY
> >
> iOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7
> C0%
> >
> 7C%7C%7C&sdata=1uUKHOonR9CtiJfdofsRbWLbuKYHlWgBXdwByQWdk2A%3
> D&reserved
> > =0
> > ---
> >   drivers/pci/controller/pci-hyperv.c | 14 +++++++++++++-
> >   1 file changed, 13 insertions(+), 1 deletion(-)
> >
> > diff --git a/drivers/pci/controller/pci-hyperv.c
> > b/drivers/pci/controller/pci-hyperv.c
> > index cfc8fa403dad..c4fba0039164 100644
> > --- a/drivers/pci/controller/pci-hyperv.c
> > +++ b/drivers/pci/controller/pci-hyperv.c
> > @@ -1040,11 +1040,18 @@ static void put_pcichild(struct hv_pci_dev
> > *hpdev)
> >
> >   /*
> >    * There is no good way to get notified from
> > vmbus_onoffer_rescind(),
> > - * so let's use polling here, since this is not a hot path.
> > + * so let's use polling here, since this is not a hot path. If
> > + * wait_for_response() has been polling for
> > + PCI_RESPONSE_HANG_TIMEOUT_SEC
> > + * without either a rescind or completion, add a periodic warning.
> >    */
> > +#define PCI_RESPONSE_HANG_TIMEOUT_SEC 300
> > +
> >   static int wait_for_response(struct hv_device *hdev,
> >   			     struct completion *comp)
> >   {
> > +	unsigned long delay =
> secs_to_jiffies(PCI_RESPONSE_HANG_TIMEOUT_SEC);
> > +	u64 timeout = get_jiffies_64() + delay;
> > +
> >   	while (true) {
> >   		if (hdev->channel->rescind) {
> >   			dev_warn_once(&hdev->device, "The device is
> gone.\n"); @@ -1053,6
> > +1060,11 @@ static int wait_for_response(struct hv_device *hdev,
> >
> >   		if (wait_for_completion_timeout(comp, HZ / 10))
> >   			break;
> > +
> > +		if (time_after64(get_jiffies_64(), timeout)) {
> > +			dev_warn(&hdev->device, "PCI stuck waiting for
> response.\n");
> > +			timeout = get_jiffies_64() + delay;
> > +		}
> >   	}
> 
> There can be some enhancements in above patch to address these problems:
> 1. Logging forever every 5 minutes in case of no completion or rescind.
> 2. If we now print warning once, not knowing if completion ever arrived.
> 3. On solving pt. 1 and 2 by adding a print for completion, one should avoid
> adding a print by default for regular timely completions.
> 
> Basically something like this:
> 
> #define PCI_RESPONSE_WARN_TIMEOUT_SEC 300
> 
> static int wait_for_response(struct hv_device *hdev,
>                   struct completion *comp) {
>      unsigned long warn_at =
>          jiffies + secs_to_jiffies(PCI_RESPONSE_WARN_TIMEOUT_SEC);
>      bool warned = false;
> 
>      while (true) {
>          if (hdev->channel->rescind) {
>              dev_warn_once(&hdev->device, "The device is gone.\n");
>              return -ENODEV;
>          }
> 
>          if (wait_for_completion_timeout(comp, HZ / 10)) {
>              if (warned || time_after_eq(jiffies, warn_at))
>                  dev_warn(&hdev->device,
>                       "PCI response received after prolonged wait.\n");
>              return 0;
>          }
> 
>          if (!warned && time_after_eq(jiffies, warn_at)) {
>              dev_warn(&hdev->device,
>                   "PCI still waiting for response.\n");
>              warned = true;
>          }
>      }
> }
> 
> Regards,
> Naman

This looks better.

Thanks,
Long

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-08-25 17:00 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-25  5:17 [PATCH] PCI: hv: Warn when wait_for_response() waits indefinitely Sahil Chandna
2026-08-25  6:46 ` Naman Jain
2026-08-25 12:00   ` Sahil Chandna
2026-08-25 17:00   ` Long Li

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox